SearcharxivSearch

arXiv subjects

Tuhin Majumder

Publications and source records attributed to Tuhin Majumder.

4 recordsLinked to original sources

Large Sample Properties of Higher Order Markov Models

We study large-sample properties of higher-order Markov chains on a finite alphabet $Σ$ when the order $m_n$ is allowed to grow with the sequence length $n$. By embedding the process into a first-order chain on $Σ^{m_n}$ and exploiting return-time decompositions, we establish a central limit theorem for additive functionals $\sum_{t}\! g_n(Y_t^{(n)})$ under natural ergodicity and sparsity conditions. The normalization involves the stationary return time to a suitably chosen state and accommodates triangular arrays with $m_n\!\to\!\infty$ and $m_n/n\!\to\!0$. We further illustrate the assumptions in a binary variable length Markov chain (VLMC), deriving explicit lower bounds on stationary masses that yield a concrete growth regime (e.g., $m_n\log m_n/n \to 0$) ensuring the CLT. These results provide asymptotic foundations for inference in sparse/partitioned higher-order models; including VLMCs and sparse Markov models (SMMs) where the effective dimensionality grows with the sample size.

math.ST

Fitting Sparse Markov Models to Categorical Time Series Using Convex Clustering

Higher-order Markov chains are frequently used to model categorical time series. However, a major problem with fitting such models is the exponentially growing number of parameters in the model order. A popular approach to parsimonious modeling is to use a Variable Length Markov Chain (VLMC), which determines relevant contexts (recent pasts) of variable orders and forms a context tree. A more general parsimonious modeling approach is given by Sparse Markov Models (SMMs), where all possible histories of order $m$ are partitioned such that the transition probability vectors are identical for the histories belonging to any particular group. In this paper, we develop an elegant method of fitting SMMs based on convex clustering and regularization. The regularization parameter is selected using the BIC criterion. Theoretical results establish model selection consistency of our method for large sample size. Extensive simulation results under different set-ups are presented to study finite sample performance of the method. Real data analysis on modelling and classifying disease sub-types demonstrates the applicability of our method as well.

stat.ME

General Robust Bayes Pseudo-Posterior: Exponential Convergence results with Applications

Although Bayesian inference is an immensely popular paradigm among a large segment of scientists including statisticians, most applications consider objective priors and need critical investigations (Efron, 2013, Science). While it has several optimal properties, a major drawback of Bayesian inference is the lack of robustness against data contamination and model misspecification, which becomes pernicious in the use of objective priors. This paper presents the general formulation of a Bayes pseudo-posterior distribution yielding robust inference. Exponential convergence results related to the new pseudo-posterior and the corresponding Bayes estimators are established under the general parametric set-up and illustrations are provided for the independent stationary as well as non-homogeneous models. Several additional details and properties of the procedure are described, including the estimation under fixed-design regression models.

math.ST

On Robust Pseudo-Bayes Estimation for the Independent Non-homogeneous Set-up

The ordinary Bayes estimator based on the posterior density suffers from the potential problems of non-robustness under data contamination or outliers. In this paper, we consider the general set-up of independent but non-homogeneous (INH) observations and study a robustified pseudo-posterior based estimation for such parametric INH models. In particular, we focus on the $R^{(α)}$-posterior developed by Ghosh and Basu (2016) for IID data and later extended by Ghosh and Basu (2017) for INH set-up, where its usefulness and desirable properties have been numerically illustrated. In this paper, we investigate the detailed theoretical properties of this robust pseudo Bayes $R^{(α)}$-posterior and associated $R^{(α)}$-Bayes estimate under the general INH set-up with applications to fixed-design regressions. We derive a Bernstein von-Mises types asymptotic normality results and Laplace type asymptotic expansion of the $R^{(α)}$-posterior as well as the asymptotic distributions of the expected $R^{(α)}$-posterior estimators. The required conditions and the asymptotic results are simplified for linear regressions with known or unknown error variance and logistic regression models with fixed covariates. The robustness of the $R^{(α)}$-posterior and associated estimators are theoretically examined through appropriate influence function analyses under general INH set-up; illustrations are provided for the case of linear regression. A high breakdown point result is derived for the expected $R^{(α)}$-posterior estimators of the location parameter under a location-scale type model. Some interesting real life data examples illustrate possible applications.

math.ST