SearcharxivSearch

arXiv subjects

Matthew Harding

Publications and source records attributed to Matthew Harding.

8 recordsLinked to original sources

Adversarial Estimation of Assortment Probabilities under Independence Structure

We consider the problem of estimating assortment probabilities, which is common in operations management applications, including product bundling, advertising, etc. Existing approaches typically model each assortment as a category and apply multinomial models to estimate the choice probabilities; while computationally convenient, these methods do not exploit independence structures in the joint distribution and may therefore be statistically inefficient when the total number of items is large. Using the representation from Bahadur (1959), we relate the sparsity of the generalized correlation coefficients to the independence structure of the binary components. We formulate the problem as estimating a high-dimensional vector of generalized correlation coefficients, together with low or moderate-dimensional nuisance parameters corresponding to the marginal probabilities. We develop a regularized adversarial estimator that attains the optimal rate under standard regularity conditions while remaining computationally feasible. The framework naturally extends to settings with covariates. We apply the proposed estimators to causal inference with multiple binary treatments and show substantial finite-sample improvements over non-adaptive methods. Numerical studies corroborate the theoretical results.

math.ST

Estimation of a Factor-Augmented Linear Model with Applications Using Student Achievement Data

In many longitudinal settings, economic theory does not guide practitioners on the type of restrictions that must be imposed to solve the rotational indeterminacy of factor-augmented linear models. We study this problem and offer several novel results on identification using internally generated instruments. We propose a new class of estimators and establish large sample results using recent developments on clustered samples and high-dimensional models. We carry out simulation studies which show that the proposed approaches improve the performance of existing methods on the estimation of unknown factors. Lastly, we consider three empirical applications using administrative data of students clustered in different subjects in elementary school, high school and college.

econ.EM

Managers versus Machines: Do Algorithms Replicate Human Intuition in Credit Ratings?

We use machine learning techniques to investigate whether it is possible to replicate the behavior of bank managers who assess the risk of commercial loans made by a large commercial US bank. Even though a typical bank already relies on an algorithmic scorecard process to evaluate risk, bank managers are given significant latitude in adjusting the risk score in order to account for other holistic factors based on their intuition and experience. We show that it is possible to find machine learning algorithms that can replicate the behavior of the bank managers. The input to the algorithms consists of a combination of standard financials and soft information available to bank managers as part of the typical loan review process. We also document the presence of significant heterogeneity in the adjustment process that can be traced to differences across managers and industries. Our results highlight the effectiveness of machine learning based analytic approaches to banking and the potential challenges to high-skill jobs in the financial sector.

econ.EM

Predicting Mortality from Credit Reports

Data on hundreds of variables related to individual consumer finance behavior (such as credit card and loan activity) is routinely collected in many countries and plays an important role in lending decisions. We postulate that the detailed nature of this data may be used to predict outcomes in seemingly unrelated domains such as individual health. We build a series of machine learning models to demonstrate that credit report data can be used to predict individual mortality. Variable groups related to credit cards and various loans, mostly unsecured loans, are shown to carry significant predictive power. Lags of these variables are also significant thus indicating that dynamics also matters. Improved mortality predictions based on consumer finance data can have important economic implications in insurance markets but may also raise privacy concerns.

econ.GN

A Panel Quantile Approach to Attrition Bias in Big Data: Evidence from a Randomized Experiment

This paper introduces a quantile regression estimator for panel data models with individual heterogeneity and attrition. The method is motivated by the fact that attrition bias is often encountered in Big Data applications. For example, many users sign-up for the latest program but few remain active users several months later, making the evaluation of such interventions inherently very challenging. Building on earlier work by Hausman and Wise (1979), we provide a simple identification strategy that leads to a two-step estimation procedure. In the first step, the coefficients of interest in the selection equation are consistently estimated using parametric or nonparametric methods. In the second step, standard panel quantile methods are employed on a subset of weighted observations. The estimator is computationally easy to implement in Big Data applications with a large number of subjects. We investigate the conditions under which the parameter estimator is asymptotically Gaussian and we carry out a series of Monte Carlo simulations to investigate the finite sample properties of the estimator. Lastly, using a simulation exercise, we apply the method to the evaluation of a recent Time-of-Day electricity pricing experiment inspired by the work of Aigner and Hausman (1980).

econ.EM

Order Determination of Large Dimensional Dynamic Factor Model

Consider the following dynamic factor model: $\mathbf{R}_t=\sum_{i=0}^q \mathbf{\Lambda}_i \mathbf{f}_{t-i}+\mathbf{e}_t,t=1,...,T$, where $\mathbf{\Lambda}_i$ is an $n\times k$ loading matrix of full rank, $\{\mathbf{f}_t\}$ are i.i.d. $k\times1$-factors, and $\mathbf{e}_t$ are independent $n\times1$ white noises. Now, assuming that $n/T\to c>0$, we want to estimate the orders $k$ and $q$ respectively. Define a random matrix $$\mathbf{\Phi}_n(\tau)=\frac{1}{2T}\sum_{j=1}^T (\mathbf{R}_j \mathbf{R}_{j+\tau}^* + \mathbf{R}_{j+\tau} \mathbf{R}_j^*),$$ where $\tau\ge 0$ is an integer. When there are no factors, the matrix $\Phi_{n}(\tau)$ reduces to $$\mathbf{M}_n(\tau) = \frac{1}{2T} \sum_{j=1}^T (\mathbf{e}_j \mathbf{e}_{j+\tau}^* + \mathbf{e}_{j+\tau} \mathbf{e}_j^*).$$ When $\tau=0$, $\mathbf{M}_n(\tau)$ reduces to the usual sample covariance matrix whose ESD tends to the well known MP law and $\mathbf{\Phi}_n(0)$ reduces to the standard spike model. Hence the number $k(q+1)$ can be estimated by the number of spiked eigenvalues of $\mathbf{\Phi}_n(0)$. To obtain separate estimates of $k$ and $q$ , we have employed the spectral analysis of $\mathbf{M}_n(\tau)$ and established the spiked model analysis for $\mathbf{\Phi}_n(\tau)$.

math.ST

Scalable Bayesian Non-Negative Tensor Factorization for Massive Count Data

We present a Bayesian non-negative tensor factorization model for count-valued tensor data, and develop scalable inference algorithms (both batch and online) for dealing with massive tensors. Our generative model can handle overdispersed counts as well as infer the rank of the decomposition. Moreover, leveraging a reparameterization of the Poisson distribution as a multinomial facilitates conjugacy in the model and enables simple and efficient Gibbs sampling and variational Bayes (VB) inference updates, with a computational cost that only depends on the number of nonzeros in the tensor. The model also provides a nice interpretability for the factors; in our model, each factor corresponds to a "topic". We develop a set of online inference algorithms that allow further scaling up the model to massive tensors, for which batch inference methods may be infeasible. We apply our framework on diverse real-world applications, such as \emph{multiway} topic modeling on a scientific publications database, analyzing a political science data set, and analyzing a massive household transactions data set.

stat.ML

Strong limit of the extreme eigenvalues of a symmetrized auto-cross covariance matrix

The auto-cross covariance matrix is defined as \[\mathbf{M}_n=\frac{1} {2T}\sum_{j=1}^T\bigl(\mathbf{e}_j\mathbf{e}_{j+\tau}^*+\mathbf{e}_{j+ \tau}\mathbf{e}_j^*\bigr),\] where $\mathbf{e}_j$'s are $n$-dimensional vectors of independent standard complex components with a common mean 0, variance $\sigma^2$, and uniformly bounded $2+\eta$th moments and $\tau$ is the lag. Jin et al. [Ann. Appl. Probab. 24 (2014) 1199-1225] has proved that the LSD of $\mathbf{M}_n$ exists uniquely and nonrandomly, and independent of $\tau$ for all $\tau\ge 1$. And in addition they gave an analytic expression of the LSD. As a continuation of Jin et al. [Ann. Appl. Probab. 24 (2014) 1199-1225], this paper proved that under the condition of uniformly bounded fourth moments, in any closed interval outside the support of the LSD, with probability 1 there will be no eigenvalues of $\mathbf{M}_n$ for all large $n$. As a consequence of the main theorem, the limits of the largest and smallest eigenvalue of $\mathbf{M}_n$ are also obtained.

math.ST