SearcharxivSearch

arXiv subjects

Jen-Chieh Teng

Publications and source records attributed to Jen-Chieh Teng.

4 recordsLinked to original sources

Unsupervised Learning Under a General Semiparametric Clusterwise Elliptical Distribution: Efficient Estimation, Optimal Clustering, and Consistent Cluster Selection

We introduce a general semiparametric clusterwise elliptical distribution to assess how latent cluster structure shapes continuous outcomes. Using a subjectwise representation, we first estimate cluster-specific mean vectors and a cluster-invariant scatter matrix by minimizing a weighted sum of squares criterion augmented with a separation penalty; we provide an initialization scheme and a computational algorithm with guaranteed convergence. This initial estimator consistently recovers the true clusters and seeds a second phase that alternates pseudo-maximum likelihood (or pseudo-maximum marginal likelihood) estimation with cluster reassignment, yielding asymptotic semiparametric efficiency and an optimal clustering that asymptotically maximizes the probability of correct membership. We also propose a semiparametric information criterion for selecting the number of clusters. Monte Carlo simulations and empirical applications demonstrate strong finite-sample performance and practical value.

stat.ME

Unsupervised Learning in a General Semiparametric Clusterwise Index Distribution Model

This study introduces a general semiparametric clusterwise index distribution model to analyze how latent clusters affect the covariate-response relationships. By employing sufficient dimension reduction to account for the effects of covariates on the cluster variable, we develop a distinct method for estimating model parameters. Building on a subjectwise representation of the underlying model, the proposed separation penalty estimation method partitions individuals and estimates cluster index coefficients. We propose a convergent algorithm for this estimation procedure and incorporate a heuristic initialization to expedite optimization. The resulting partition estimator is subsequently used to fit the cluster membership model and to construct an optimal classification rule, with both procedures iteratively updating the partition and parameter estimators. Another key contribution of our method is the development of two consistent semiparametric information criteria for selecting the number of clusters. In line with principles of classification and estimation in supervised learning, the estimated cluster structure is consistent and optimal, and the parameter estimators possess the oracle property. Comprehensive simulation studies and empirical data analyses illustrate the effectiveness of the proposed methodology.

stat.ME

An Effective Method for Identifying Clusters of Robot Strengths

In the analysis of qualification data from the FIRST Robotics Competition, the ratio of the number of observations to the number of parameters has been found to be quite small for the commonly used winning margin power rating (WMPR) model. This usually leads to imprecise estimates and inaccurate predictions in such a three-on-three game. With the finding of a clustering feature in estimated robot strengths, a more flexible model with latent clusters of robots was proposed to alleviate overparameterization of the WMPR model. Since its structure can be regarded as a dimension reduction of the parameter space in the WMPR model, the identification of clusters of robot strengths is naturally transformed into a model selection problem. Instead of comparing a huge number of competing models, we develop an effective method to estimate the number of clusters, clusters of robots, and robot strengths. The new method consists of two parts: (i) a combination of hierarchical and non-hierarchical classifications to determine candidate models; and (ii) variant goodness-of-fit criteria to select optimal models. Different from existing hierarchical classification systems, each step of ours is based on estimated robot strengths from a candidate model in the preceding non-hierarchical classification step. A great advantage of the designed non-hierarchical classification system is to examine the possibility of reassigning robots to other cluster sets of robots. To reduce the overestimation of clusters by the mean squared prediction error criteria, the corresponding BIC are established as alternatives for model selection. By assembling these essential elements into a coherent whole, a systematic procedure is presented to perform the estimation. In addition, we propose two indices to measure the nested relation between cluster sets of two models and monotonic association between robot strengths of two models.

stat.AP

Estimating Robot Strengths with Application to Selection of Alliance Members in FIRST Robotics Competitions

Since the inception of the FIRST Robotics Competition (FRC) and its special playoff system, robotics teams have longed to appropriately quantify the strengths of their designed robots. The FRC includes a playground draft-like phase (alliance selection), arguably the most game-changing part of the competition, in which the top-8 robotics teams in a tournament based on the FRC's ranking system assess potential alliance members for the opportunity of partnering in a playoff stage. In such a three-versus-three competition, several measures and models have been used to characterize actual or relative robot strengths. However, existing models are found to have poor predictive performance due to their imprecise estimates of robot strengths caused by a small ratio of the number of observations to the number of robots. A more general regression model with latent clusters of robot strengths is, thus, proposed to enhance their predictive capacities. Two effective estimation procedures are further developed to simultaneously estimate the number of clusters, clusters of robots, and robot strengths. Meanwhile, some measures are used to assess the predictive ability of competing models, the agreement between published FRC measures of strength and model-based robot strengths of all, playoff, and FRC top-8 robots, and the agreement between FRC top-8 robots and model-based top robots. Moreover, the stability of estimated robot strengths and accuracies is investigated to determine whether the scheduled matches are excessive or insufficient. In the analysis of qualification data from the 2018 FRC Houston and Detroit championships, the predictive ability of our model is also shown to be significantly better than those of existing models. Teams who adopt the new model can now appropriately rank their preferences for playoff alliance partners with greater predictive capability than before.

stat.AP