SearcharxivSearch

arXiv subjects

Ting-Li Chen

Publications and source records attributed to Ting-Li Chen.

14 recordsLinked to original sources

Structure-Preserving Visualization of Complex Systems through Discrete Approximation: An Application to Argo Data

This paper presents a framework for constructing structure-preserving representations of complex systems through discrete approximation, and demonstrates its use in studying the vertical temperature and salinity structures in the mesopelagic zone across the global ocean using the ARGO dataset. Clustering serves as a means of organizing complexity into a finite set of structures that approximate the overall oceanic conditions, and a color encoding design then integrates these structures into a coherent map, with the three color components derived from interpretable geometric features of a profile: its initial level, its magnitude of variation, and its shape. Instead of focusing on specific depth levels or computing zonal averages within selected regions, our approach preserves the full vertical structure of individual profiles and incorporates each profile in the global ocean, capturing both fine-scale profile detail and large-scale spatial variability. By clustering over one million profiles collected over a decade, we identify and characterize representative profile shapes, which form the basis for a visualization strategy that provides an integrated, comprehensive, and interpretable presentation of the large-scale spatial distributions of these oceanic vertical patterns.

stat.ME

Simultaneous Estimation of Ballpark Effects and Team Defense Using Total Bases Residuals

Estimating ballpark effects and team defense in baseball is challenging because batted-ball outcomes are influenced by multiple factors, including contact quality, ballpark environment, defensive performance, and random variation. In this study, we propose a simple and interpretable framework based on Total Bases Residuals (TBR). Using Statcast data from 2015 to 2024, we construct expected total bases conditional on exit velocity and launch angle, and define residuals relative to this baseline. These residuals allow us to separate the effects of ballpark environment and team defense and to estimate them simultaneously within a unified regression framework. Our results show that, when our estimates differ from official MLB metrics, they are more consistent with observed home-away patterns for both teams and their opponents, providing empirical support for our approach. Similar patterns are also observed in comparisons with existing defensive metrics. The results also suggest changes in league-wide outcomes and are broadly consistent with developments in the game, including the increased use of data-driven positioning, the restriction on defensive shifts, and possible changes in the physical properties of the baseball. We further introduce a standardized index that facilitates comparison across teams, ballparks, and seasons by expressing effects in units of standard deviation.

stat.AP

Blurring Mean Shift for Clustering Functional Data: A Scalable Algorithm and Convergence Analysis

This paper extends the blurring mean shift algorithm from vector-valued data to functional data, enabling effective clustering in infinite-dimensional settings without requiring specification of the number of clusters. To address the computational challenges posed by large-scale datasets, we introduce a fast stochastic variant that significantly reduces computational complexity. We provide a rigorous convergence analysis for the full blurring functional mean shift procedure, establishing theoretical guarantees for its iterative behavior. For the stochastic variant, we provide partial theoretical justification by showing that, when the subset size is sufficiently large, its one-step update is well approximated by the corresponding update of the full algorithm. The proposed method is demonstrated through real-data applications, including hourly Taiwan PM$_{2.5}$ measurements and Argo oceanographic profiles. Our key contributions include: (1) extending the blurring mean shift algorithm to functional data in a Hilbert-space setting; (2) developing a scalable stochastic variant based on random partitioning for large-scale data; (3) establishing convergence results for the full blurring functional mean shift algorithm; and (4) demonstrating the scalability and practical usefulness of the proposed method through simulation and real-data applications.

stat.ME

Unraveling implicit human behavioral effects on dynamic characteristics of Covid-19 daily infection rates in Taiwan

We study Covid-19 spreading dynamics underlying 84 curves of daily Covid-19 infection rates pertaining to 84 districts belonging to the largest seven cities in Taiwan during her pristine surge period. Our computational developments begin with selecting and extracting 18 features from each smoothed district-specific curve. This step of computing effort allows unstructured data to be converted into structured data, with which we then demonstrate asymmetric growth and decline dynamics among all involved curves. Specifically, based on Theoretical Information measurements of conditional entropy and mutual information, we compute major factors of order-1 and order-2 that reveal significant effects on affecting the curves' peak value and curvature at peak, which are two essential features characterizing all the curves. Further, we investigate and demonstrate major factors determining the geographic and social-economic induced behavioral effects by encoding each of these 84 districts with two binary characteristics: North-vs-South and Unban-vs-suburban. Furthermore, based on this data-driven knowledge on the district scale, we go on to study fine-scale behavioral effects on infectious disease spreading through similarity among 96 age-group-specific curves of daily infection rate within 12 urban districts of Taipei and 12 suburban districts of New Taipei City, which counts for almost one-quarter of the island nation's total population. We conclude that human living, traveling, and working behaviors do implicitly affect the spreading dynamics of Covid-19 across Taiwan profoundly.

stat.AP

Multiscale major factor selections for complex system data with structural dependency and heterogeneity

Based on structured data derived from large complex systems, we computationally further develop and refine a major factor selection protocol by accommodating structural dependency and heterogeneity among many features to unravel data's information content. Two operational concepts: ``de-associating'' and its counterpart ``shadowing'' that play key roles in our protocol, are reasoned, explained, and carried out via contingency table platforms. This protocol via ``de-associating'' capability would manifest data's information content by identifying which covariate feature-sets do or don't provide information beyond the first identified major factors to join the collection of major factors as secondary members. Our computational developments begin with globally characterizing a complex system by structural dependency between multiple response (Re) features and many covariate (Co) features. We first apply our major factor selection protocol on a Behavioral Risk Factor Surveillance System (BRFSS) data set to demonstrate discoveries of localities where heart-diseased patients become either majorities or further reduced minorities that sharply contrast data's imbalance nature. We then study a Major League Baseball (MLB) data set consisting of 12 pitchers across 3 seasons, reveal detailed multiscale information content regarding pitching dynamics, and provide nearly perfect resolutions to the Multiclass Classification (MCC) problem and the difficult task of detecting idiosyncratic changes of any individual pitcher across multiple seasons. We conclude by postulating an intuitive conjecture that large complex systems related to inferential topics can only be efficiently resolved through discoveries of data's multiscale information content reflecting the system's authentic structural dependency and heterogeneity.

stat.ME

Learned practical guidelines for evaluating Conditional Entropy and Mutual Information in discovering major factors of response-vs-covariate dynamics

We reformulate and reframe a series of increasingly complex parametric statistical topics into a framework of response-vs-covariate (Re-Co) dynamics that is described without any explicit functional structures. Then we resolve these topics' data analysis tasks by discovering major factors underlying such Re-Co dynamics by only making use of data's categorical nature. The major factor selection protocol at the heart of Categorical Exploratory Data Analysis (CEDA) paradigm is illustrated and carried out by employing Shannon's conditional entropy (CE) and mutual information ($I[Re; Co] $) as two key Information Theoretical measurements. Through the process of evaluating these two entropy-based measurements and resolving statistical tasks, we acquire several computational guidelines for carrying out the major factor selection protocol in a do-and-learn fashion. Specifically, practical guidelines are established for evaluating CE and $I[Re; Co] $ in accord with the criterion called [C1:confirmable]. Via [C1:confirmable] criterion, we make no attempts on acquiring consistent estimations of these theoretical information measurements. All evaluations are carried out on a contingency table platform, upon which the practical guidelines also provide ways of lessening effects of curse of dimensionality. We explicitly carry out six examples of Re-Co dynamics, within each of which, several widely extended scenarios are also explored and discussed.

stat.ME

A Consistency Theorem for Randomized Singular Value Decomposition

The singular value decomposition (SVD) and the principal component analysis are fundamental tools and probably the most popular methods for data dimension reduction. The rapid growth in the size of data matrices has lead to a need for developing efficient large-scale SVD algorithms. Randomized SVD was proposed, and its potential was demonstrated for computing a low-rank SVD (Rokhlin et al., 2009). In this article, we provide a consistency theorem for the randomized SVD algorithm and a numerical example to show how the random projections to low dimension affect the consistency.

math.ST

On the asymptotic variance of reversible Markov chain without cycles

Markov chain Monte Carlo(MCMC) is a popular approach to sample from high dimensional distributions, and the asymptotic variance is a commonly used criterion to evaluate the performance. While most popular MCMC algorithms are reversible, there is a growing literature on the development and analyses of nonreversible MCMC. Chen and Hwang(2013) showed that a reversible MCMC can be improved by adding an antisymmetric perturbation. They also raised a conjecture that it can not be improved if there is no cycle in the corresponding graph. In this paper, we present a rigorous proof of this conjecture. The proof is based on the fact that the transition matrix with an acyclic structure will produce minimum commute time between vertices.

math.PR

Integrating multiple random sketches for singular value decomposition

The singular value decomposition (SVD) of large-scale matrices is a key tool in data analytics and scientific computing. The rapid growth in the size of matrices further increases the need for developing efficient large-scale SVD algorithms. Randomized SVD based on one-time sketching has been studied, and its potential has been demonstrated for computing a low-rank SVD. Instead of exploring different single random sketching techniques, we propose a Monte Carlo type integrated SVD algorithm based on multiple random sketches. The proposed integration algorithm takes multiple random sketches and then integrates the results obtained from the multiple sketched subspaces. So that the integrated SVD can achieve higher accuracy and lower stochastic variations. The main component of the integration is an optimization problem with a matrix Stiefel manifold constraint. The optimization problem is solved using Kolmogorov-Nagumo-type averages. Our theoretical analyses show that the singular vectors can be induced by population averaging and ensure the consistencies between the computed and true subspaces and singular vectors. Statistical analysis further proves a strong Law of Large Numbers and gives a rate of convergence by the Central Limit Theorem. Preliminary numerical results suggest that the proposed integrated SVD algorithm is promising.

math.NA

On the strengths of the self-updating process clustering algorithm

We introduce a simple, intuitive and yet powerful algorithm for clustering analysis. This algorithm is an iterative process on the sample space, which arises as an extension of the iteratively generated correlation matrices. It allows for both time-varying and time-invariant operators, therefore can be considered more general than the blurring mean-shift algorithm in which operators are time-invariant. The algorithm stands from the viewpoint of data points and simulates the process how data points move and perform self-clustering, therefore is named Self-Updating Process (SUP). It is particularly competitive for (i) data with noise, (ii) data with large number of clusters and (iii) unbalanced data. When noise is present in the data, the algorithm is able to isolate noisy points while performing clustering simultaneously. Simulation studies and real data applications are presented to demonstrate the performance of SUP.

stat.ME

Functional Inverse Regression in an Enlarged Dimension Reduction Space

We consider an enlarged dimension reduction space in functional inverse regression. Our operator and functional analysis based approach facilitates a compact and rigorous formulation of the functional inverse regression problem. It also enables us to expand the possible space where the dimension reduction functions belong. Our formulation provides a unified framework so that the classical notions, such as covariance standardization, Mahalanobis distance, SIR and linear discriminant analysis, can be naturally and smoothly carried out in our enlarged space. This enlarged dimension reduction space also links to the linear discriminant space of Gaussian measures on a separable Hilbert space.

math.ST

On the Weak Convergence and Central Limit Theorem of Blurring and Nonblurring Processes with Application to Robust Location Estimation

This article studies the weak convergence and associated Central Limit Theorem for blurring and nonblurring processes. Then, they are applied to the estimation of location parameter. Simulation studies show that the location estimation based on the convergence point of blurring process is more robust and often more efficient than that of nonblurring process.

math.ST

$γ$-SUP: A clustering algorithm for cryo-electron microscopy images of asymmetric particles

Cryo-electron microscopy (cryo-EM) has recently emerged as a powerful tool for obtaining three-dimensional (3D) structures of biological macromolecules in native states. A minimum cryo-EM image data set for deriving a meaningful reconstruction is comprised of thousands of randomly orientated projections of identical particles photographed with a small number of electrons. The computation of 3D structure from 2D projections requires clustering, which aims to enhance the signal to noise ratio in each view by grouping similarly oriented images. Nevertheless, the prevailing clustering techniques are often compromised by three characteristics of cryo-EM data: high noise content, high dimensionality and large number of clusters. Moreover, since clustering requires registering images of similar orientation into the same pixel coordinates by 2D alignment, it is desired that the clustering algorithm can label misaligned images as outliers. Herein, we introduce a clustering algorithm $γ$-SUP to model the data with a $q$-Gaussian mixture and adopt the minimum $γ$-divergence for estimation, and then use a self-updating procedure to obtain the numerical solution. We apply $γ$-SUP to the cryo-EM images of two benchmark macromolecules, RNA polymerase II and ribosome. In the former case, simulated images were chosen to decouple clustering from alignment to demonstrate $γ$-SUP is more robust to misalignment outliers than the existing clustering methods used in the cryo-EM community. In the latter case, the clustering of real cryo-EM data by our $γ$-SUP method eliminates noise in many views to reveal true structure features of ribosome at the projection level.

stat.AP

On the Convergence and Consistency of the Blurring Mean-Shift Process

The mean-shift algorithm is a popular algorithm in computer vision and image processing. It can also be cast as a minimum gamma-divergence estimation. In this paper we focus on the "blurring" mean shift algorithm, which is one version of the mean-shift process that successively blurs the dataset. The analysis of the blurring mean-shift is relatively more complicated compared to the nonblurring version, yet the algorithm convergence and the estimation consistency have not been well studied in the literature. In this paper we prove both the convergence and the consistency of the blurring mean-shift. We also perform simulation studies to compare the efficiency of the blurring and the nonblurring versions of the mean-shift algorithms. Our results show that the blurring mean-shift has more efficiency.

stat.ML