SearcharxivSearch

arXiv subjects

Tzee-Ming Huang

Publications and source records attributed to Tzee-Ming Huang.

7 recordsLinked to original sources

Testing for cross-quantilogram change

For two time series $\{ (Y_t, Z_t^Y) \}_{t}$ and $\{(X_t, Z_t^X)\}_{t}$, the directional dependence of $\{ X_t \}_{t}$ on $\{ Y_t \}_{t}$ while removing the impact of $Z_t^X$ on $X_t$ and the impact of $Z_t^Y$ on $ Y_t$ can be measured by cross-quantilograms. When the two time series are obeserved over two periods of time, it can be of interest to learn whether the cross-quantilograms remain the same for the two periods of time. We propose a test for this purpose, and the cross-quantilograms are estimated using the estimators proposed by Han (2016). The $p$-value of the proposed test is obtained based on a bootstrap approach.

stat.ME

A mixture logistic model for panel data with a Markov structure

In this study, we propose a mixture logistic regression model with a Markov structure, and consider the estimation of model parameters using maximum likelihood estimation. We also provide a forward type variable selection algorithm to choose the important explanatory variables to reduce the number of parameters in the proposed model.

stat.ME

Random Partitioning and Distribution-based Thresholding for Iterative Variable Screening in High Dimensions

In big data analysis, a simple task such as linear regression can become very challenging as the variable dimension $p$ grows. As a result, variable screening is inevitable in many scientific studies. In recent years, randomized algorithms have become a new trend and are playing an increasingly important role for large scale data analysis. In this article, we combine the ideas of variable screening and random partitioning to propose a new iterative variable screening method. For moderate sized $p$ of order $O(n^{2-δ})$, we propose a basic algorithm that adopts a distribution-based thresholding rule. For very large $p$, we further propose a two-stage procedure. This two-stage procedure first performs a random partitioning to divide predictors into subsets of manageable size of order $O(n^{2-δ})$ for variable screening, where $δ>0$ can be an arbitrarily small positive number. Random partitioning is repeated a few times. Next, the final estimate of variable subset is obtained by integrating results obtained from multiple random partitions. Simulation studies show that our method works well and outperforms some renowned competitors. Real data applications are presented. Our algorithms are able to handle predictors in the size of millions.

stat.ME

A clustering method for misaligned curves

We consider the problem of clustering misaligned curves. According to our similarity measure, two curves are considered similar if they have the same shape after being aligned, and the warping function does not differ from the identity function very much. A clustering method is proposed, which updates curves so that similar curves become more similar, and then combines curves that are similar enough to form clusters. The proposed method needs to be used together with a clustering index and a set of combination thresholds. Simulation results are presented to demonstrate the performance of this approach under different parameter settings and clustering indexes. Two real data applications are included.

stat.ME

A nonparametric copula density estimator incorporating information on bivariate marginals

We propose a copula density estimator that can include information on bivariate marginals when the information is available. We use B-splines for copula density approximation and include information on bivariate marginals via a penalty term. Our estimator satisfies the constraints for a copula density. Under mild conditions, the proposed estimator is consistent.

stat.ME

Testing conditional independence using maximal nonlinear conditional correlation

In this paper, the maximal nonlinear conditional correlation of two random vectors $X$ and $Y$ given another random vector $Z$, denoted by $ρ_1(X,Y|Z)$, is defined as a measure of conditional association, which satisfies certain desirable properties. When $Z$ is continuous, a test for testing the conditional independence of $X$ and $Y$ given $Z$ is constructed based on the estimator of a weighted average of the form $\sum_{k=1}^{n_Z}f_Z(z_k)ρ^2_1(X,Y|Z=z_k)$, where $f_Z$ is the probability density function of $Z$ and the $z_k$'s are some points in the range of $Z$. Under some conditions, it is shown that the test statistic is asymptotically normal under conditional independence, and the test is consistent.

math.ST

Convergence rates for posterior distributions and adaptive estimation

The goal of this paper is to provide theorems on convergence rates of posterior distributions that can be applied to obtain good convergence rates in the context of density estimation as well as regression. We show how to choose priors so that the posterior distributions converge at the optimal rate without prior knowledge of the degree of smoothness of the density function or the regression function to be estimated.

math.ST