SearcharxivSearch

arXiv subjects

Daniel A. Griffith

Publications and source records attributed to Daniel A. Griffith.

9 recordsLinked to original sources

Some Simplifications for the Expectation-Maximization (EM) Algorithm: The Linear Regression Model Case

The EM algorithm is a generic tool that offers maximum likelihood solutions when datasets are incomplete with data values missing at random or completely at random. At least for its simplest form, the algorithm can be rewritten in terms of an ANCOVA regression specification. This formulation allows several analytical results to be derived that permit the EM algorithm solution to be expressed in terms of new observation predictions and their variances. Implementations can be made with a linear regression or a nonlinear regression model routine, allowing missing value imputations, even when they must satisfy constraints. Fourteen example datasets gleaned from the EM algorithm literature are reanalyzed. Imputation results have been verified with SAS PROC MI. Six theorems are proved that broadly contextualize imputation findings in terms of the theory, methodology, and practice of statistical science.

stat.ME

Sub-model aggregation for scalable eigenvector spatial filtering: Application to spatially varying coefficient modeling

This study proposes a method for aggregating/synthesizing global and local sub-models for fast and flexible spatial regression modeling. Eigenvector spatial filtering (ESF) was used to model spatially varying coefficients and spatial dependence in the residuals by sub-model, while the generalized product-of-experts method was used to aggregate these sub-models. The major advantages of the proposed method are as follows: (i) it is highly scalable for large samples in terms of accuracy and computational efficiency; (ii) it is easily implemented by estimating sub-models independently first and aggregating/averaging them thereafter; and (iii) likelihood-based inference is available because the marginal likelihood is available in closed-form. The accuracy and computational efficiency of the proposed method are confirmed using Monte Carlo simulation experiments. This method was then applied to residential land price analysis in Japan. The results demonstrate the usefulness of this method for improving the interpretability of spatially varying coefficients. The proposed method is implemented in an R package spmoran (version 0.3.0 or later).

stat.ME

Balancing spatial and non-spatial variation in varying coefficient modeling: a remedy for spurious correlation

This study discusses the importance of balancing spatial and non-spatial variation in spatial regression modeling. Unlike spatially varying coefficients (SVC) modeling, which is popular in spatial statistics, non-spatially varying coefficients (NVC) modeling has largely been unexplored in spatial fields. Nevertheless, as we will explain, consideration of non-spatial variation is needed not only to improve model accuracy but also to reduce spurious correlation among varying coefficients, which is a major problem in SVC modeling. We consider a Moran eigenvector approach modeling spatially and non-spatially varying coefficients (S&NVC). A Monte Carlo simulation experiment comparing our S&NVC model with existing SVC models suggests both modeling accuracy and computational efficiency for our approach. Beyond that, somewhat surprisingly, our approach identifies true and spurious correlations among coefficients nearly perfectly, even when usual SVC models suffer from severe spurious correlations. It implies that S&NVC model should be used even when the analysis purpose is modeling SVCs. Finally, our S&NVC model is employed to analyze a residential land price dataset. Its results suggest existence of both spatial and non-spatial variation in regression coefficients in practice. The S&NVC model is now implemented in the R package spmoran.

stat.AP

A memory-free spatial additive mixed modeling for big spatial data

This study develops a spatial additive mixed modeling (AMM) approach estimating spatial and non-spatial effects from large samples, such as millions of observations. Although fast AMM approaches are already well-established, they are restrictive in that they assume an known spatial dependence structure. To overcome this limitation, this study develops a fast AMM with the estimation of spatial structure in residuals and regression coefficients together with non-spatial effects. We rely on a Moran coefficient-based approach to estimate the spatial structure. The proposed approach pre-compresses large matrices whose size grows with respect to the sample size N before the model estimation; thus, the computational complexity for the estimation is independent of the sample size. Furthermore, the pre-compression is done through a block-wise procedure that makes the memory consumption independent of N. Eventually, the spatial AMM is memory-free and fast even for millions of observations. The developed approach is compared to alternatives through Monte Carlo simulation experiments. The result confirms the accuracy and computational efficiency of the developed approach. The developed approaches are implemented in an R package spmoran.

stat.ME

Low rank spatial econometric models

This article presents a re-structuring of spatial econometric models in a linear mixed model framework. To that end, it proposes low rank spatial econometric models that are robust to the existence of noise (i.e., measurement error), and can enjoy fast parameter estimation and inference by Type II restricted likelihood maximization (empirical Bayes) techniques. The small sample properties of the proposed low rank spatial econometric models are examined using Monte Carlo simulation experiments, the results of these experiments confirm that direct effects and indirect effects a la LeSage and Pace (2009) can be estimated with a high degree of accuracy. Also, when data are noisy, estimators for coefficients in the proposed models have lower root mean squared errors compared to conventional specifications, despite them being low rank approximations. The proposed approach is implemented in an R package "spmoran".

stat.ME

Spatially varying coefficient modeling for large datasets: Eliminating N from spatial regressions

While spatially varying coefficient (SVC) modeling is popular in applied science, its computational burden is substantial. This is especially true if a multiscale property of SVC is considered. Given this background, this study develops a Moran's eigenvector-based spatially varying coefficients (M-SVC) modeling approach that estimates multiscale SVCs computationally efficiently. This estimation is accelerated through a (i) rank reduction, (ii) pre-compression, and (iii) sequential likelihood maximization. Steps (i) and (ii) eliminate the sample size N from the likelihood function; after these steps, the likelihood maximization cost is independent of N. Step (iii) further accelerates the likelihood maximization so that multiscale SVCs can be estimated even if the number of SVCs, K, is large. The M-SVC approach is compared with geographically weighted regression (GWR) through Monte Carlo simulation experiments. These simulation results show that our approach is far faster than GWR when N is large, despite numerically estimating 2K parameters while GWR numerically estimates only 1 parameter. Then, the proposed approach is applied to a land price analysis as an illustration. The developed SVC estimation approach is implemented in the R package "spmoran."

stat.ME

Eigenvector spatial filtering for large data sets: fixed and random effects approaches

Eigenvector spatial filtering (ESF) is a spatial modeling approach, which has been applied in urban and regional studies, ecological studies, and so on. However, it is computationally demanding, and may not be suitable for large data modeling. The objective of this study is developing fast ESF and random effects ESF (RE-ESF), which are capable of handling very large samples. To achieve it, we accelerate eigen-decomposition and parameter estimation, which make ESF and RE-ESF slow. The former is accelerated by utilizing the Nystrom extension, whereas the latter is by small matrix tricks. The resulting fast ESF and fast RE-ESF are compared with non-approximated ESF and RE-ESF in Monte Carlo simulation experiments. The result shows that, while ESF and RE-ESF are slow for several thousand samples, fast ESF and RE-ESF require only several seconds for the samples. They also suggest that the proposed approaches effectively remove positive spatial dependence in the residuals with very small approximation errors when the number of eigenvectors considered is 200 or more. Note that these approaches cannot deal with negative spatial dependence. The proposed approaches are implemented in an R package "spmoran."

stat.ME

The importance of scale in spatially varying coefficient modeling

While spatially varying coefficient (SVC) models have attracted considerable attention in applied science, they have been criticized as being unstable. The objective of this study is to show that capturing the "spatial scale" of each data relationship is crucially important to make SVC modeling more stable, and in doing so, adds flexibility. Here, the analytical properties of six SVC models are summarized in terms of their characterization of scale. Models are examined through a series of Monte Carlo simulation experiments to assess the extent to which spatial scale influences model stability and the accuracy of their SVC estimates. The following models are studied: (i) geographically weighted regression (GWR) with a fixed distance or (ii) an adaptive distance bandwidth (GWRa), (iii) flexible bandwidth GWR (FB-GWR) with fixed distance or (iv) adaptive distance bandwidths (FB-GWRa), (v) eigenvector spatial filtering (ESF), and (vi) random effects ESF (RE-ESF). Results reveal that the SVC models designed to capture scale dependencies in local relationships (FB-GWR, FB-GWRa and RE-ESF) most accurately estimate the simulated SVCs, where RE-ESF is the most computationally efficient. Conversely GWR and ESF, where SVC estimates are naively assumed to operate at the same spatial scale for each relationship, perform poorly. Results also confirm that the adaptive bandwidth GWR models (GWRa and FB-GWRa) are superior to their fixed bandwidth counterparts (GWR and FB-GWR).

stat.ME

A Moran coefficient-based mixed effects approach to investigate spatially varying relationships

This study develops a spatially varying coefficient model by extending the random effects eigenvector spatial filtering model. The developed model has the following properties: its coefficients are interpretable in terms of the Moran coefficient; each of its coefficients can have a different degree of spatial smoothness; and it yields a variant of a Bayesian spatially varying coefficient model. Also, parameter estimation of the model can be executed with a relatively small computationally burden. Results of a Monte Carlo simulation reveal that our model outperforms a conventional eigenvector spatial filtering (ESF) model and geographically weighted regression (GWR) models in terms of the accuracy of the coefficient estimates and computational time. We empirically apply our model to the hedonic land price analysis of flood risk in Japan.

stat.ME