SearcharxivSearch

arXiv subjects

Fushing Hsieh

Publications and source records attributed to Fushing Hsieh.

15 recordsLinked to original sources

Design of Experiment in Complex Systems based on Computational Taxonomy

Via Computational Taxonomy (CT), we develop Design of Experiment(DoE) based on rigorously redefined constituting ingredients of complex system dynamics: randomness, nonlinearity and even class, through a data-driven constructed Taxonomic Hierarchy. As an opposite quest of Classification without man-made assumptions and structures, we illustrate this new perspective of DoE through a civil engineering complex system: Concrete Compressive Strength (CCS). CT begins by building a Taxonomic Hierarchy as a heterogeneity-vs-homogeneity map framed with a tree geometry to represent CCS-system dynamics. At each internal node of this hierarchy, a heatmap is computed via Scientific Data Analysis (SDA) to reveal locality-embraced heterogeneity through block structured covariate homogeneity annotated with response's locality-split. Only arriving at each ending-node of this hierarchy, coherence of homogeneity is achieved on both response and covariate sides. As such a class of finite sample nature is computationally recognized and confirmed. In contrast, nonlinearity is evidently observed as incoherence of response-vs-covariate homogeneity when comparing two classes located on two distinct branches. This hierarchy explicitly maps out system's randomness and nonlinearity to serve as a scientific basis for any DoE quest. Design of Experiment (DoE) for any designated class is redefined as a search for a covariate subspace that embraces the class-representative randomness and at the same time avoids potential nonlinearities with respect to the rest of classes. This is a brand-new theme of DoE.

stat.ME

Structure-Preserving Visualization of Complex Systems through Discrete Approximation: An Application to Argo Data

This paper presents a framework for constructing structure-preserving representations of complex systems through discrete approximation, and demonstrates its use in studying the vertical temperature and salinity structures in the mesopelagic zone across the global ocean using the ARGO dataset. Clustering serves as a means of organizing complexity into a finite set of structures that approximate the overall oceanic conditions, and a color encoding design then integrates these structures into a coherent map, with the three color components derived from interpretable geometric features of a profile: its initial level, its magnitude of variation, and its shape. Instead of focusing on specific depth levels or computing zonal averages within selected regions, our approach preserves the full vertical structure of individual profiles and incorporates each profile in the global ocean, capturing both fine-scale profile detail and large-scale spatial variability. By clustering over one million profiles collected over a decade, we identify and characterize representative profile shapes, which form the basis for a visualization strategy that provides an integrated, comprehensive, and interpretable presentation of the large-scale spatial distributions of these oceanic vertical patterns.

stat.ME

Unraveling heterogeneity of ADNI's time-to-event data using conditional entropy Part-I: Cross-sectional study

Through Alzheimer's Disease Neuroimaging Initiative (ADNI), time-to-event data: from the pre-dementia state of mild cognitive impairment (MCI) to the diagnosis of Alzheimer's disease (AD), is collected and analyzed by explicitly unraveling prognostic heterogeneity among 346 uncensored and 557 right censored subjects under structural dependency among covariate features. The non-informative censoring mechanism is tested and confirmed based on conditional-vs-marginal entropies evaluated upon contingency tables built by the Redistribute-to-the-right algorithm. The Categorical Exploratory Data Analysis (CEDA) paradigm is applied to evaluate conditional entropy-based associative patterns between the categorized response variable against 16 categorized covariable variables all having 4 categories. Two order-1 global major factors: V9 (MEM-mean) and V8 (ADAS13.bl) are selected sharing the highest amounts of mutual information with the response variable. This heavily censored data set is analyzed by Cox's proportional hazard (PH) modeling. Comparisons of PH and CEDA results on a global scale are complicated under the structural dependency of covariate features. To alleviate such complications, V9 and V8 are taken as two potential perspectives of heterogeneity and the entire collections of subjects are divided into two sets of four sub-collections. CEDA major factor selection protocol is applied to all sub-collections to figure out which features provide extra information. Graphic displays are developed to explicitly unravel conditional entropy expansions upon perspectives of heterogeneity in ADNI data. On the local scale, PH analysis is carried out and results are compared with CEDA's. We conclude that, when facing structural dependency among covariates and heterogeneity in data, CEDA and its major factor selection provide significant merits for manifesting data's multiscale information content.

stat.AP

An Encoding Approach for Stable Change Point Detection

Without imposing prior distributional knowledge underlying multivariate time series of interest, we propose a nonparametric change-point detection approach to estimate the number of change points and their locations along the temporal axis. We develop a structural subsampling procedure such that the observations are encoded into multiple sequences of Bernoulli variables. A maximum likelihood approach in conjunction with a newly developed searching algorithm is implemented to detect change points on each Bernoulli process separately. Then, aggregation statistics are proposed to collectively synthesize change-point results from all individual univariate time series into consistent and stable location estimations. We also study a weighting strategy to measure the degree of relevance for different subsampled groups. Simulation studies are conducted and shown that the proposed change-point methodology for multivariate time series has favorable performance comparing with currently popular nonparametric methods under various settings with different degrees of complexity. Real data analyses are finally performed on categorical, ordinal, and continuous time series taken from fields of genetics, climate, and finance.

stat.ME

Coarse- and fine-scale geometric information content of Multiclass Classification and implied Data-driven Intelligence

Under any Multiclass Classification (MCC) setting defined by a collection of labeled point-cloud specified by a feature-set, we extract only stochastic partial orderings from all possible triplets of point-cloud without explicitly measuring the three cloud-to-cloud distances. We demonstrate that such a collective of partial ordering can efficiently compute a label embedding tree geometry on the Label-space. This tree in turn gives rise to a predictive graph, or a network with precisely weighted linkages. Such two multiscale geometries are taken as the coarse scale information content of MCC. They indeed jointly shed lights on explainable knowledge on why and how labeling comes about and facilitates error-free prediction with potential multiple candidate labels supported by data. For revealing within-label heterogeneity, we further undergo labeling naturally found clusters within each point-cloud, and likewise derive multiscale geometry as its fine-scale information content contained in data. This fine-scale endeavor shows that our computational proposal is indeed scalable to a MCC setting having a large label-space. Overall the computed multiscale collective of data-driven patterns and knowledge will serve as a basis for constructing visible and explainable subject matter intelligence regarding the system of interest.

stat.ML

Discovering Multiple Phases of Dynamics by Dissecting Multivariate Time Series

We proposed a data-driven approach to dissect multivariate time series in order to discover multiple phases underlying dynamics of complex systems. This computing approach is developed as a multiple-dimension version of Hierarchical Factor Segmentation(HFS) technique. This expanded approach proposes a systematic protocol of choosing various extreme events in multi-dimensional space. Upon each chosen event, an empirical distribution of event-recurrence, or waiting time between the excursions, is fitted by a geometric distribution with time-varying parameters. Iterative fittings are performed across all chosen events. We then collect and summarize the local recurrent patterns into a global dynamic mechanism. Clustering is applied for partitioning the whole time period into alternating segments, in which variables are identically distributed. Feature weighting techniques are also considered to compensate for some drawbacks of clustering. Our simulation results show that this expanded approach can even detect systematic differences when the joint distribution varies. In real data experiments, we analyze the relationship from returns, trading volume, and transaction number of a single, as well as of multiple stocks in S&P500. We can successfully not only map out volatile periods but also provide potential associative links between stocks.

stat.ME

Unraveling S&P500 stock volatility and networks -- An encoding-and-decoding approach

Volatility of financial stock is referring to the degree of uncertainty or risk embedded within a stock's dynamics. Such risk has been received huge amounts of attention from diverse financial researchers. By following the concept of regime-switching model, we proposed a non-parametric approach, named encoding-and-decoding, to discover multiple volatility states embedded within a discrete time series of stock returns. The encoding is performed across the entire span of temporal time points for relatively extreme events with respect to a chosen quantile-based threshold. As such the return time series is transformed into Bernoulli-variable processes. In the decoding phase, we computationally seek for locations of change points via estimations based on a new searching algorithm in conjunction with the information criterion applied on the observed collection of recurrence times upon the binary process. Besides the independence required for building the Geometric distributional likelihood function, the proposed approach can functionally partition the entire return time series into a collection of homogeneous segments without any assumptions of dynamic structure and underlying distributions. In the numerical experiments, our approach is found favorably compared with parametric models like Hidden Markov Model. In the real data applications, we introduce the application of our approach in forecasting stock returns. Finally, volatility dynamic of every single stock of S&P500 is revealed, and a stock network is consequently established to represent dependency relations derived through concurrent volatility states among S&P500.

q-fin.ST

Categorical exploratory data analysis on goodness-of-fit issues

If the aphorism "All models are wrong"- George Box, continues to be true in data analysis, particularly when analyzing real-world data, then we should annotate this wisdom with visible and explainable data-driven patterns. Such annotations can critically shed invaluable light on validity as well as limitations of statistical modeling as a data analysis approach. In an effort to avoid holding our real data to potentially unattainable or even unrealistic theoretical structures, we propose to utilize the data analysis paradigm called Categorical Exploratory Data Analysis (CEDA). We illustrate the merits of this proposal with two real-world data sets from the perspective of goodness-of-fit. In both data sets, the Normal distribution's bell shape seemingly fits rather well by first glance. We apply CEDA to bring out where and how each data fits or deviates from the model shape via several important distributional aspects. We also demonstrate that CEDA affords a version of tree-based p-value, and compare it with p-values based on traditional statistical approaches. Along our data analysis, we invest computational efforts in making graphic display to illuminate the advantages of using CEDA as one primary way of data analysis in Data Science education.

stat.ML

Extreme-K categorical samples problem

With histograms as its foundation, we develop Categorical Exploratory Data Analysis (CEDA) under the extreme-$K$ sample problem, and illustrate its universal applicability through four 1D categorical datasets. Given a sizable $K$, CEDA's ultimate goal amounts to discover by data's information content via carrying out two data-driven computational tasks: 1) establish a tree geometry upon $K$ populations as a platform for discovering a wide spectrum of patterns among populations; 2) evaluate each geometric pattern's reliability. In CEDA developments, each population gives rise to a row vector of categories proportions. Upon the data matrix's row-axis, we discuss the pros and cons of Euclidean distance against its weighted version for building a binary clustering tree geometry. The criterion of choice rests on degrees of uniformness in column-blocks framed by this binary clustering tree. Each tree-leaf (population) is then encoded with a binary code sequence, so is tree-based pattern. For evaluating reliability, we adopt row-wise multinomial randomness to generate an ensemble of matrix mimicries, so an ensemble of mimicked binary trees. Reliability of any observed pattern is its recurrence rate within the tree ensemble. A high reliability value means a deterministic pattern. Our four applications of CEDA illuminate four significant aspects of extreme-$K$ sample problems.

stat.AP

Color-complexity enabled exhaustive color-dots identification and spatial patterns testing in images

Targeted color-dots with varying shapes and sizes in images are first exhaustively identified, and then their multiscale 2D geometric patterns are extracted for testing spatial uniformness in a progressive fashion. Based on color theory in physics, we develop a new color-identification algorithm relying on highly associative relations among the three color-coordinates: RGB or HSV. Such high associations critically imply low color-complexity of a color image, and renders potentials of exhaustive identification of targeted color-dots of all shapes and sizes. Via heterogeneous shaded regions and lighting conditions, our algorithm is shown being robust, practical and efficient comparing with the popular Contour and OpenCV approaches. Upon all identified color-pixels, we form color-dots as individually connected networks with shapes and sizes. We construct minimum spanning trees (MST) as spatial geometries of dot-collectives of various size-scales. Given a size-scale, the distribution of distances between immediate neighbors in the observed MST is extracted, so do many simulated MSTs under the spatial uniformness assumption. We devise a new algorithm for testing 2D spatial uniformness based on a Hierarchical clustering tree upon all involving MSTs. Our developments are illustrated on images obtained by mimicking chemical spraying via drone in Precision Agriculture.

cs.CV

Categorical Exploratory Data Analysis: From Multiclass Classification and Response Manifold Analytics perspectives of baseball pitching dynamics

From two coupled Multiclass Classification (MCC) and Response Manifold Analytics (RMA) perspectives, we develop Categorical Exploratory Data Analysis (CEDA) on PITCHf/x database for the information content of Major League Baseball's (MLB) pitching dynamics. MCC and RMA information contents are represented by one collection of multi-scales pattern categories from mixing geometries and one collection of global-to-local geometric localities from response-covariate manifolds, respectively. These collectives shed light on the pitching dynamics and maps out uncertainty of popular machine learning approaches. On MCC setting, an indirect-distance-measure based label embedding tree leads to discover asymmetry of mixing geometries among labels' point-clouds. A selected chain of complementary covariate feature groups collectively brings out multi-order mixing geometric pattern categories. Such categories then reveal the true nature of MCC predictive inferences. On RMA setting, multiple response features couple with multiple major covariate features to demonstrate physical principles bearing manifolds with a lattice of natural localities. With minor features' heterogeneous effects being locally identified, such localities jointly weave their focal characteristics into system understanding and provide a platform for RMA predictive inferences. Our CEDA works for universal data types, adopts non-linear associations and facilitates efficient feature-selections and inferences.

stat.AP

From learning gait signatures of many individuals to reconstructing gait dynamics of one single individual

Based on the same databases, we computationally address two seemingly highly related, in fact drastically distinct, questions via computational data-driven algorithms: 1) how to precisely achieve the big task of differentiating gait signatures of many individuals? 2) how to reconstruct an individual's complex gait dynamics in full? Our brains can "effortlessly" resolve the first question, but will definitely fail in the second one. Since many fine temporal scale gait patterns surely escape our eyes. Based on accelerometers' 3D gait time series databases, we link the answers toward both questions via multiscale structural dependency within gait dynamics of our musculoskeletal system. Two types of dependency manifestations are explored. We first develop simple algorithmic computing called Principle System-State Analysis (PSSA) for the coarse dependency in implicit forms. PSSA is shown to be able to efficiently classifying among many subjects. We then develop a multiscale Local-1st-Global-2nd (L1G2) Coding Algorithm and a landmark computing algorithm. With both algorithms, we can precisely dissect rhythmic gait cycles, and then decompose each cycle into a series of cyclic gait phases. With proper color-coding and stacking, we reconstruct and represent an individual's gait dynamics via a 3D cylinder to collectively reveal universal deterministic and stochastic structural patterns on centisecond (10 milliseconds) scale across all rhythmic cycles. This 3D cylinder can serve as "passtensor" for authentication purposes related to clinical diagnoses and cybersecurity.

eess.SP

Information of Epileptic Mechanism and its Systemic Change-points in a Zebrafish's Brain-wide Calcium Imaging Video Data

The epileptic mechanism is postulated as that an animal's neurons gradually diminish their inhibition function coupled with enhanced excitation when an epileptic event is approaching. Calcium imaging technique is designed to directly record brain-wide neurons activity in order to discover the underlying epileptic mechanism. In this paper, using one brain-wide calcium imaging video of Zebrafish, we compute dynamic pattern information of the epileptic mechanism, and devise three graphical displays to show the visible functional aspect of epileptic mechanism over five inter-ictal periods. The foundation of our data-driven computations for such dynamic patterns relies on one universal phenomenon discovered across 696 informative pixels. This universality is that each pixel's progressive 5-percentile process oscillates in an irregular fashion at first, but, after the middle point of inter-ictal period, the oscillation is replaced by a steady increasing trend. Such dynamic patterns are collectively transformed into a visible systemic change-point as an early warning signal (EWS) of an incoming epileptic event. We conclude through the graphic displays that pattern information extracted from the calcium imaging video realistically reveals the Zebrafish's authentic epileptic mechanism.

q-bio.NC

Graphic displays of MLB pitching mechanics and its evolutions in PITCHf/x data

Systemic and idiosyncratic patterns in pitching mechanics of 24 top starting pitchers in Major League Baseball (MLB) are extracted and discovered from PITCHf/x database. These evolving patterns across different pitchers or seasons are represented through three exclusively developed graphic displays. Understanding on such patterned evolutions will be beneficial for pitchers' wellbeing in signaling potential injury, and will be critical for expert knowledge in comparing pitchers. Based on data-driven computing, a universal composition of patterns is identified on all pitchers' mutual conditional entropy matrices. The first graphic display reveals that this universality accommodates physical laws as well as systemic characteristics of pitching mechanics. Such visible characters point to large scale factors for differentiating between distinct clusters of pitchers, and simultaneously lead to detailed factors for comparing individual pitchers. The second graphic display shows choices of features that are able to express a pitcher's season-by-season pitching contents via a series of 3(+2)D point-cloud geometries. The third graphic display exhibits exquisitely a pitcher's idiosyncratic pattern-information of pitching across seasons by demonstrating all his pitch-subtype evolutions. These heatmap-based graphic displays are platforms for visualizing and understanding pitching mechanics.

stat.AP

Complexity of Possibly-gapped Histogram and Analysis of Histogram (ANOHT)

Without unrealistic continuity and smoothness assumptions on a distributional density of one dimensional dataset, constructing an authentic possibly-gapped histogram becomes rather complex. The candidate ensemble is described via a two-layer Ising model, and its size is shown to grow exponentially. This exponential complexity makes any exhaustive search in-feasible and all boundary parameters local. For data compression via Uniformity, the decoding error criterion is nearly independent of sample size. These characteristics nullify statistical model selection techniques, such as Minimum Description Length (MDL). Nonetheless practical and nearly optimal solutions are algorithmically computable. A data-driven algorithm is devised to construct such histograms along the branching hierarchy of a Hierarchical Clustering tree. Such resultant histograms naturally manifest data's physical information contents: deterministic structures of bin-boundaries coupled with stochastic structures of Uniformity within each bin. Without enforcing unrealistic Normality and constant variance assumptions, an application of possibly-gapped histogram is devised, called analysis of Histogram (ANOHT), to replace Analysis of Variance (ANOVA). Its potential applications are foreseen in digital re-normalization schemes and associative pattern extraction among features of heterogeneous data types. Thus constructing possibly-gapped histograms becomes a prerequisite for knowledge discovery, via exploratory data analysis and unsupervised Machine Learning.

stat.ME