SearcharxivSearch

arXiv subjects

Semhar Michael

Publications and source records attributed to Semhar Michael.

5 recordsLinked to original sources

A reduced rank model for spatial categorical data with many classes

We develop an identifiable reduced-rank spatial multinomial model for categorical data with many classes. The model represents class-specific spatial effects through a low-dimensional set of shared latent factors, substantially reducing parameter dimension while preserving joint dependence across classes. Because standard conjugate and P\'olya-Gamma methods fail under this factorization, we propose a Gibbs sampler using Laplace-approximation proposals within Metropolis-Hastings updates. Simulation studies examine dimension selection and the accuracy of the Laplace proposals. An application to dominant tree species mapping in the Blue Ridge Mountains demonstrates scalable inference and flexible joint predictions for individual classes, class unions, and area-level summaries.

stat.ME

Estimation of Parameters of the Truncated Normal Distribution with Unknown Bounds

Estimators of parameters of truncated distributions, namely the truncated normal distribution, have been widely studied for a known truncation region. There is also literature for estimating the unknown bounds for known parent distributions. In this work, we develop a novel algorithm under the expectation-solution (ES) framework, which is an iterative method of solving nonlinear estimating equations, to estimate both the bounds and the location and scale parameters of the parent normal distribution utilizing the theory of best linear unbiased estimates from location-scale families of distribution and unbiased minimum variance estimation of truncation regions. The conditions for the algorithm to converge to the solution of the estimating equations for a fixed sample size are discussed, and the asymptotic properties of the estimators are characterized using results on M- and Z-estimation from empirical process theory. The proposed method is then compared to methods utilizing the known truncation bounds via Monte Carlo simulation.

stat.CO

Statistical few-shot learning for large-scale classification via parameter pooling

In large-scale few-shot learning for classification problems, often there are a large number of classes and few high-dimensional observations per class. Previous model-based methods, such as Fisher's linear discriminant analysis (LDA), require the strong assumptions of a shared covariance matrix between all classes. Quadratic discriminant analysis will often lead to singular or unstable covariance matrix estimates. Both of these methods can lead to lower-than-desired classification performance. We introduce a novel, model-based clustering method that can relax the shared covariance assumptions of LDA by clustering sample covariance matrices, either singular or non-singular. In addition, we study the statistical properties of parameter estimates. This will lead to covariance matrix estimates which are pooled within each cluster of classes. We show, using simulated and real data, that our classification method tends to yield better discrimination compared to other methods.

stat.ME

Learning trends of COVID-19 using semi-supervised clustering

A finite mixture model is used to learn trends from the currently available data on coronavirus (COVID-19). Data on the number of confirmed COVID-19 related cases and deaths for European countries and the United States (US) are explored. A semi-supervised clustering approach with positive equivalence constraints is used to incorporate country and state information into the model. The analysis of trends in the rates of cases and deaths is carried out jointly using a mixture of multivariate Gaussian non-linear regression models with a mean trend specified using a generalized logistic function. The optimal number of clusters is chosen using the Bayesian information criterion. The resulting clusters provide insight into different mitigation strategies adopted by US states and European countries. The obtained results help identify the current relative standing of individual states and show a possible future if they continue with the chosen mitigation technique

stat.AP

Exploring the Daschle Collection using Text Mining

A U.S. Senator from South Dakota donated documents that were accumulated during his service as a house representative and senator to be housed at the Bridges library at South Dakota State University. This project investigated the utility of quantitative statistical methods to explore some portions of this vast document collection. The available scanned documents and emails from constituents are analyzed using natural language processing methods including the Latent Dirichlet Allocation (LDA) model. This model identified major topics being discussed in a given collection of documents. Important events and popular issues from the Senator Daschles career are reflected in the changing topics from the model. These quantitative statistical methods provide a summary of the massive amount of text without requiring significant human effort or time and can be applied to similar collections.

cs.IR