SearcharxivSearch

arXiv subjects

Emanuel Ben-David

Publications and source records attributed to Emanuel Ben-David.

14 recordsLinked to original sources

Spatial Dependence in the Self-Response: Spatial Dependence, Modeling, and Operational Consequences

The U.S.\ Census Bureau's Low Response Score (LRS) is a central planning instrument for identifying places likely to require additional self-response outreach and nonresponse follow-up. The published LRS is intentionally interpretable: it is built from tract-level covariates using an ordinary least squares specification. That transparency, however, leaves open an important question for official statistics: how much spatial structure remains after the own-tract covariates have done their work, and what form does that structure take? Using the observed 2010 Census mail non-return rate for 71,076 U.S. census tracts and the twenty-five Erdman--Bates LRS predictors, this paper compares the full spatial autoregressive model family under queen-contiguity weights and validates the leading candidates with both random and spatial-block cross-validation. OLS leaves strong residual spatial autocorrelation ($I=0.399$). Formal diagnostics and model comparisons indicate that the remaining dependence is primarily error-type rather than a global endogenous lag process. Although the spatial Durbin model minimizes in-sample AIC, spatial-block validation reverses that ranking: the error-family models (SEM/SDEM) generalize best, while the AIC-best SDM is weakest out of sample. The SDEM provides an interpretable middle ground, absorbing residual spatial dependence while representing neighborhood demographic effects as local spillovers. Robustness checks show that these conclusions are invariant to the weights definition and are not an artifact of tract-size-driven heteroskedasticity. The results suggest that LRS-style response models should be evaluated with spatial validation, not only in-sample fit, and that local neighborhood context can be operationally meaningful without invoking a global response-contagion mechanism.

stat.AP

A General Framework for Regression with Mismatched Data Based on Mixture Modeling

Data sets obtained from linking multiple files are frequently affected by mismatch error, as a result of non-unique or noisy identifiers used during record linkage. Accounting for such mismatch error in downstream analysis performed on the linked file is critical to ensure valid statistical inference. In this paper, we present a general framework to enable valid post-linkage inference in the challenging secondary analysis setting in which only the linked file is given. The proposed framework covers a wide selection of statistical models and can flexibly incorporate additional information about the underlying record linkage process. Specifically, we propose a mixture model for pairs of linked records whose two components reflect distributions conditional on match status, i.e., correct match or mismatch. Regarding inference, we develop a method based on composite likelihood and the EM algorithm as well as an extension towards a fully Bayesian approach. Extensive simulations and several case studies involving contemporary record linkage applications corroborate the effectiveness of our framework.

stat.ME

Regularization for Shuffled Data Problems via Exponential Family Priors on the Permutation Group

In the analysis of data sets consisting of (X, Y)-pairs, a tacit assumption is that each pair corresponds to the same observation unit. If, however, such pairs are obtained via record linkage of two files, this assumption can be violated as a result of mismatch error rooting, for example, in the lack of reliable identifiers in the two files. Recently, there has been a surge of interest in this setting under the term "Shuffled data" in which the underlying correct pairing of (X, Y)-pairs is represented via an unknown index permutation. Explicit modeling of the permutation tends to be associated with substantial overfitting, prompting the need for suitable methods of regularization. In this paper, we propose a flexible exponential family prior on the permutation group for this purpose that can be used to integrate various structures such as sparse and locally constrained shuffling. This prior turns out to be conjugate for canonical shuffled data problems in which the likelihood conditional on a fixed permutation can be expressed as product over the corresponding (X,Y)-pairs. Inference is based on the EM algorithm in which the intractable E-step is approximated by the Fisher-Yates algorithm. The M-step is shown to admit a significant reduction from $n^2$ to $n$ terms if the likelihood of (X,Y)-pairs has exponential family form as in the case of generalized linear models. Comparisons on synthetic and real data show that the proposed approach compares favorably to competing methods.

stat.ML

Predicting Census Survey Response Rates With Parsimonious Additive Models and Structured Interactions

In this paper, we consider the problem of predicting survey response rates using a family of flexible and interpretable nonparametric models. The study is motivated by the US Census Bureau's well-known ROAM application, which uses a linear regression model trained on the US Census Planning Database data to identify hard-to-survey areas. A crowdsourcing competition (Erdman and Bates, 2016) organized more than ten years ago revealed that machine learning methods based on ensembles of regression trees led to the best performance in predicting survey response rates; however, the corresponding models could not be adopted for the intended application due to their black-box nature. We consider nonparametric additive models with a small number of main and pairwise interaction effects using $\ell_0$-based penalization. From a methodological viewpoint, we study our estimator's computational and statistical aspects and discuss variants incorporating strong hierarchical interactions. Our algorithms (open-sourced on GitHub) extend the computational frontiers of existing algorithms for sparse additive models to be able to handle datasets relevant to the application we consider. We discuss and interpret findings from our model on the US Census Planning Database. In addition to being useful from an interpretability standpoint, our models lead to predictions comparable to popular black-box machine learning methods based on gradient boosting and feedforward neural networks - suggesting that it is possible to have models that have the best of both worlds: good model accuracy and interpretability.

stat.ML

Estimation in exponential family Regression based on linked data contaminated by mismatch error

Identification of matching records in multiple files can be a challenging and error-prone task. Linkage error can considerably affect subsequent statistical analysis based on the resulting linked file. Several recent papers have studied post-linkage linear regression analysis with the response variable in one file and the covariates in a second file from the perspective of the "Broken Sample Problem" and "Permuted Data". In this paper, we present an extension of this line of research to exponential family response given the assumption of a small to moderate number of mismatches. A method based on observation-specific offsets to account for potential mismatches and $\ell_1$-penalization is proposed, and its statistical properties are discussed. We also present sufficient conditions for the recovery of the correct correspondence between covariates and responses if the regression parameter is known. The proposed approach is compared to established baselines, namely the methods by Lahiri-Larsen and Chambers, both theoretically and empirically based on synthetic and real data. The results indicate that substantial improvements over those methods can be achieved even if only limited information about the linkage process is available.

stat.ME

A Two-Stage Approach to Multivariate Linear Regression with Sparsely Mismatched Data

A tacit assumption in linear regression is that (response, predictor)-pairs correspond to identical observational units. A series of recent works have studied scenarios in which this assumption is violated under terms such as ``Unlabeled Sensing and ``Regression with Unknown Permutation''. In this paper, we study the setup of multiple response variables and a notion of mismatches that generalizes permutations in order to allow for missing matches as well as for one-to-many matches. A two-stage method is proposed under the assumption that most pairs are correctly matched. In the first stage, the regression parameter is estimated by handling mismatches as contaminations, and subsequently the generalized permutation is estimated by a basic variant of matching. The approach is both computationally convenient and equipped with favorable statistical guarantees. Specifically, it is shown that the conditions for permutation recovery become considerably less stringent as the number of responses $m$ per observation increase. Particularly, for $m = Ω(\log n)$, the required signal-to-noise ratio no longer depends on the sample size $n$. Numerical results on synthetic and real data are presented to support the main findings of our analysis.

stat.ML

A Pseudo-Likelihood Approach to Linear Regression with Partially Shuffled Data

Recently, there has been significant interest in linear regression in the situation where predictors and responses are not observed in matching pairs corresponding to the same statistical unit as a consequence of separate data collection and uncertainty in data integration. Mismatched pairs can considerably impact the model fit and disrupt the estimation of regression parameters. In this paper, we present a method to adjust for such mismatches under ``partial shuffling" in which a sufficiently large fraction of (predictors, response)-pairs are observed in their correct correspondence. The proposed approach is based on a pseudo-likelihood in which each term takes the form of a two-component mixture density. Expectation-Maximization schemes are proposed for optimization, which (i) scale favorably in the number of samples, and (ii) achieve excellent statistical performance relative to an oracle that has access to the correct pairings as certified by simulations and case studies. In particular, the proposed approach can tolerate considerably larger fraction of mismatches than existing approaches, and enables estimation of the noise level as well as the fraction of mismatches. Inference for the resulting estimator (standard errors, confidence intervals) can be based on established theory for composite likelihood estimation. Along the way, we also propose a statistical test for the presence of mismatches and establish its consistency under suitable conditions.

stat.ME

Linear Regression with Sparsely Permuted Data

In regression analysis of multivariate data, it is tacitly assumed that response and predictor variables in each observed response-predictor pair correspond to the same entity or unit. In this paper, we consider the situation of "permuted data" in which this basic correspondence has been lost. Several recent papers have considered this situation without further assumptions on the underlying permutation. In applications, the latter is often to known to have additional structure that can be leveraged. Specifically, we herein consider the common scenario of "sparsely permuted data" in which only a small fraction of the data is affected by a mismatch between response and predictors. However, an adverse effect already observed for sparsely permuted data is that the least squares estimator as well as other estimators not accounting for such partial mismatch are inconsistent. One approach studied in detail herein is to treat permuted data as outliers which motivates the use of robust regression formulations to estimate the regression parameter. The resulting estimate can subsequently be used to recover the permutation. A notable benefit of the proposed approach is its computational simplicity given the general lack of procedures for the above problem that are both statistically sound and computationally appealing.

math.ST

Sharper lower and upper bounds for the gaussian rank of a graph

An open problem in graphical Gaussian models is to determine the smallest number of observations needed to guarantee the existence of the maximum likelihood estimator of the covariance matrix with probability one. In this paper we formalize a closely related problem in which the existence of the maximum likelihood estimator is guaranteed for all generic observations. We call the number determined by this problem the Gaussian rank of the graph representing the model. We prove that the Gaussian rank is strictly between the subgraph connectivity number and the graph degeneracy number. These bounds are in general much sharper than the best bounds known in the literature and furthermore computable in polynomial time.

math.ST

The Letac-Massam conjecture and existence of high dimensional Bayes estimators for Graphical Models

In recent years, a variety of useful extensions of the Wishart have been proposed in the literature for the purposes of studying Markov random fields/graphical models. In particular, generalizations of the Wishart, referred to as Type I and Type II Wishart distributions, have been introduced by Letac and Massam (\emph{Annals of Statistics} 2006) and play important roles in both frequentist and Bayesian inference for Gaussian graphical models. These distributions have been especially useful in high-dimensional settings due to the flexibility offered by their multiple shape parameters. The domain of In this paper we resolve a long-standing conjecture of Letac and Massam (LM) concerning the domains of the multi-parameters of graphical Wishart type distributions. This conjecture, posed in \emph{Annals of Statistics}, also relates fundamentally to the existence of Bayes estimators corresponding to these high dimensional priors. To achieve our goal, we first develop novel theory in the context of probabilistic analysis of graphical models. Using these tools, and a recently introduced class of Wishart distributions for directed acyclic graph (DAG) models, we proceed to give counterexamples to the LM conjecture, thus completely resolving the problem. Our analysis also proceeds to give useful insights on graphical Wishart distributions with implications for Bayesian inference for such models.

math.ST

Positive definite completion problems for directed acyclic graphs

A positive definite completion problem pertains to determining whether the unspecified positions of a partial (or incomplete) matrix can be completed in a desired subclass of positive definite matrices. In this paper we study an important and new class of positive definite completion problems where the desired subclasses are the spaces of covariance and inverse-covariance matrices of probabilistic models corresponding to directed acyclic graph models (also known as Bayesian networks). We provide fast procedures that determine whether a partial matrix can be completed in either of these spaces and thereafter proceed to construct the completed matrices. We prove an analog of the positive definite completion result for undirected graphs in the context of directed acyclic graphs, and thus proceed to characterize the class of DAGs which can always be completed. We also proceed to give closed form expressions for the inverse and the determinant of a completed matrix as a function of only the elements of the corresponding partial matrix.

math.RA

Maximal Invariants Over Symmetric Cones

In this paper we consider some hypothesis tests within a family of Wishart distributions, where both the sample space and the parameter space are symmetric cones. For such testing problems, we first derive the joint density of the ordered eigenvalues of the generalized Wishart distribution and propose a test statistic analog to that of classical multivariate statistics for testing homoscedasticity of covariance matrix. In this generalization of Bartlett's test for equality of variances to hypotheses of real, complex, quaternion, Lorentz and octonion types of covariance structures.

math.ST

Maximal Invariants For Lorentz Wishart Models

In this paper we consider two statistical hypotheses for the families of Wishart type distributions. These distributions are analogs of the Wishart distributions defined and parametrized over a Lorentz cone. We test these hypotheses by means of maximal invariant statistics which are explicitly derived in the paper. The testing problems, respectively, concern the hypothesis that parameters are in a sub-Lorentz-cone, and the the hypothesis that two observations have the same parameter.

math.ST

High dimensional Bayesian inference for Gaussian directed acyclic graph models

We study centered Gaussian models Markov with respect to a directed acyclic graph (DAG) whose vertices have a fixed parent ordering. We construct a conjugate family on the modified Cholesky parameters, with one shape parameter per vertex, and derive its induced distributions on incomplete covariance and precision coordinates. The distribution is proper exactly when \(\alpha_i>|\mathrm{pa}(i)|+2\) for every vertex, and its nodewise conditional-variance and regression parameters are independent across vertices. This factorization gives a closed-form normalizing constant, conjugate updating, marginal likelihoods, and explicit full-matrix posterior means. We distinguish these posterior means from nonlinear completions of incomplete-coordinate means. We also distinguish the transformed Cholesky-coordinate mode from modes defined using covariance or precision coordinates. The model-selection procedure searches only over DAGs compatible with the specified ordering. Historical simulation and data examples illustrate the method; their evidentiary limitations and reproducibility requirements are stated explicitly.

math.ST