SearcharxivSearch

arXiv subjects

Fan Dai

Publications and source records attributed to Fan Dai.

10 recordsLinked to original sources

Multi-layered model-based characterisation of the local-Universe galaxy data from the GAMA survey

Understanding the formation and evolution of galaxy populations requires robust classification and characterisation techniques that jointly account for internal galaxy properties and environment. We analyse $5,306$ galaxies from the Galaxy And Mass Assembly (GAMA) survey, described by stellar mass, specific star formation rate, $u-r$ colour, half-light radius, S\'ersic index, and a combined environmental measure given by the optimal density. Unlike distance-based unsupervised clustering methods, our framework provides a probabilistic characterisation of galaxy populations, accommodates heavy-tailed feature distributions, and captures dependence among observables through latent factors. We model the sample using a $t$-mixture of factor analysers with group-specific latent structures (M$t$FAD), and then apply model-estimated overlap-based syncytial clustering (MOBSynC) to merge weakly separated groups and recover higher-level population structure. The first stage identifies eight simple clusters. The third and the fourth groups lie on the red, low-star-forming sequence and correspond to environmentally quenched and mass-quenched systems, respectively, while the sixth group traces the massive end of the star-forming sequence, and the seventh group appears to represent a more heterogeneous population that may include transition objects. The remaining groups populate the low- to intermediate-mass blue sequence, including both compact and more extended star-forming galaxies. The second MOBSynC stage merges the simple clusters into two compound groups: a red sequence formed by the third and the fourth groups, and the rest merging to form a broad blue sequence. Our results show that the familiar red-blue bimodality of local galaxies contains additional physically meaningful substructure linked to quenching pathway, morphology, and environment.

astro-ph.GA

Boxplots and quartile plots for grouped and periodic angular data

Angular observations, or observations lying on the unit circle, arise in many disciplines and require special care in their description, analysis, interpretation and visualization. We provide methods to construct concentric circular boxplot displays of distributions of groups of angular data. The use of concentric boxplots brings challenges of visual perception, so we set the boxwidths to be inversely proportional to the square root of their distance from the centre. A perception survey supports this scaled boxwidth choice. For a large number of groups, we propose circular quartile plots. A three-dimensional toroidal display is also implemented for periodic angular distributions. We illustrate our methods on datasets in (1) psychology, to display motor resonance under different conditions, (2) genomics, to understand the distribution of peak phases for ancillary clock genes, and (3) meteorology and wind turbine power generation, to study the changing and periodic distribution of wind direction over the course of a year.

stat.ME

Latent characterisation of the complete BATSE gamma ray bursts catalogue using Gaussian mixture of factor analysers and model-estimated overlap-based syncytial clustering

Characterising and distinguishing gamma-ray bursts (GRBs) has interested astronomers for many decades. While some authors have found two or three groups of GRBs by analyzing only a few parameters, recent work identified five ellipsoidally-shaped groups upon considering nine parameters $T_{50}, T_{90}, F_1, F_2, F_3, F_4, P_{64}, P_{256}, P_{1024}$. Yet others suggest sub-classes within the two or three groups found earlier. Using a mixture model of Gaussian factor analysers, we analysed 1150 GRBs, that had nine parameters observed, from the current Burst and Transient Source Experiment (BATSE) catalogue, and again established five ellipsoidal-shaped groups to describe the GRBs. These five groups are characterised in terms of their average duration, fluence and spectrum as shorter-faint-hard, long-intermediate-soft, long-intermediate-intermediate, long-bright-intermediate and short-faint-hard. The use of factor analysers in describing individual group densities allows for a more thorough group-wise characterisation of the parameters in terms of a few latent features. However, given the discrepancy with many other existing studies that advocated for two or three groups, we also performed model-estimated overlap-based syncytial clustering (MOBSynC) that successively merges poorer-separated groups. The five ellipsoidal groups merge into three and then into two groups, one with GRBs of low durations and the other having longer duration GRBs. These groups are also characterised in terms of a few latent factors made up of the nine parameters. Our analysis provides context for all three sets of results, and in doing so, details a multi-layered characterisation of the BATSE GRBs, while also explaining the structure in their variability.

astro-ph.HE

Democratizing planetary-scale analysis: An ultra-lightweight Earth embedding database for accurate and flexible global land monitoring

The rapid evolution of satellite-borne Earth Observation (EO) systems has revolutionized terrestrial monitoring, yielding petabyte-scale archives. However, the immense computational and storage requirements for global-scale analysis often preclude widespread use, hindering planetary-scale studies. To address these barriers, we present Embedded Seamless Data (ESD), an ultra-lightweight, 30-m global Earth embedding database spanning the 25-year period from 2000 to 2024. By transforming high-dimensional, multi-sensor observations from the Landsat series (5, 7, 8, and 9) and MODIS Terra into information-dense, quantized latent vectors, ESD distills essential geophysical and semantic features into a unified latent space. Utilizing the ESDNet architecture and Finite Scalar Quantization (FSQ), the dataset achieves a transformative ~340-fold reduction in data volume compared to raw archives. This compression allows the entire global land surface for a single year to be encapsulated within approximately 2.4 TB, enabling decadal-scale global analysis on standard local workstations. Rigorous validation demonstrates high reconstructive fidelity (MAE: 0.0130; RMSE: 0.0179; CC: 0.8543). By condensing the annual phenological cycle into 12 temporal steps, the embeddings provide inherent denoising and a semantically organized space that outperforms raw reflectance in land-cover classification, achieving 79.74% accuracy (vs. 76.92% for raw fusion). With robust few-shot learning capabilities and longitudinal consistency, ESD provides a versatile foundation for democratizing planetary-scale research and advancing next-generation geospatial artificial intelligence.

cs.CV

A Hybrid Mixture Approach for Clustering and Characterizing Cancer Data

Model-based clustering is widely used for identifying and distinguishing types of diseases. However, modern biomedical data coming with high dimensions make it challenging to perform the model estimation in traditional cluster analysis. The incorporation of factor analyzer into the mixture model provides a way to characterize the large set of data features, but the current estimation method is computationally impractical for massive data due to the intrinsic slow convergence of the embedded algorithms, and the incapability to vary the size of the factor analyzers, preventing the implementation of a generalized mixture of factor analyzers and further characterization of the data clusters. We propose a hybrid matrix-free computational scheme to efficiently estimate the clusters and model parameters based on a Gaussian mixture along with generalized factor analyzers to summarize the large number of variables using a small set of underlying factors. Our approach outperforms the existing method with faster convergence while maintaining high clustering accuracy. Our algorithms are applied to accurately identify and distinguish types of breast cancer based on large tumor samples, and to provide a generalized characterization for subtypes of lymphoma using massive gene records.

stat.ME

A Hybrid Mixture of $t$-Factor Analyzers for Clustering High-dimensional Data

This paper develops a novel hybrid approach for estimating the mixture model of $t$-factor analyzers (MtFA) that employs multivariate $t$-distribution and factor model to cluster and characterize grouped data. The traditional estimation method for MtFA faces computational challenges, particularly in high-dimensional settings, where the eigendecomposition of large covariance matrices and the iterative nature of Expectation-Maximization (EM) algorithms lead to scalability issues. We propose a computational scheme that integrates a profile likelihood method into the EM framework to efficiently obtain the model parameter estimates. The effectiveness of our approach is demonstrated through simulations showcasing its superior computational efficiency compared to the existing method, while preserving clustering accuracy and resilience against outliers. Our method is applied to cluster the Gamma-ray bursts, reinforcing several claims in the literature that Gamma-ray bursts have heterogeneous subpopulations and providing characterizations of the estimated groups.

stat.ME

Exploratory Factor Analysis of Data on a Sphere

Data on high-dimensional spheres arise frequently in many disciplines either naturally or as a consequence of preliminary processing and can have intricate dependence structure that needs to be understood. We develop exploratory factor analysis of the projected normal distribution to explain the variability in such data using a few easily interpreted latent factors. Our methodology provides maximum likelihood estimates through a novel fast alternating expectation profile conditional maximization algorithm. Results on simulation experiments on a wide range of settings are uniformly excellent. Our methodology provides interpretable and insightful results when applied to tweets with the $\#MeToo$ hashtag in early December 2018, to time-course functional Magnetic Resonance Images of the average pre-teen brain at rest, to characterize handwritten digits, and to gene expression data from cancerous cells in the Cancer Genome Atlas.

stat.ME

Fully Three-dimensional Radial Visualization

We develop methodology for three-dimensional (3D) radial visualization (RadViz) of multidimensional datasets. The classical two-dimensional (2D) RadViz visualizes multivariate data in the 2D plane by mapping every observation to a point inside the unit circle. Our tool, RadViz3D, distributes anchor points uniformly on the 3D unit sphere. We show that this uniform distribution provides the best visualization with minimal artificial visual correlation for data with uncorrelated variables. However, anchor points can be placed exactly equi-distant from each other only for the five Platonic solids, so we provide equi-distant anchor points for these five settings, and approximately equi-distant anchor points via a Fibonacci grid for the other cases. Our methodology, implemented in the R package $radviz3d$, makes fully 3D RadViz possible and is shown to improve the ability of this nonlinear technique in more faithfully displaying simulated data as well as the crabs, olive oils and wine datasets. Additionally, because radial visualization is naturally suited for compositional data, we use RadViz3D to illustrate (i) the chemical composition of Longquan celadon ceramics and their Jingdezhen imitation over centuries, and (ii) US regional SARS-Cov-2 variants' prevalence in the Covid-19 pandemic during the summer 2021 surge of the Delta variant.

stat.ME

A Matrix--free Likelihood Method for Exploratory Factor Analysis of High-dimensional Gaussian Data

This paper proposes a novel profile likelihood method for estimating the covariance parameters in exploratory factor analysis of high-dimensional Gaussian datasets with fewer observations than number of variables. An implicitly restarted Lanczos algorithm and a limited-memory quasi-Newton method are implemented to develop a matrix-free framework for likelihood maximization. Simulation results show that our method is substantially faster than the expectation-maximization solution without sacrificing accuracy. Our method is applied to fit factor models on data from suicide attempters, suicide ideators and a control group.

stat.ME

Visualization of Labeled Mixed-featured Datasets

We develop methodology for visualization of labeled mixed-featured datasets. We first investigate datasets with continuous features where our Max-Ratio Projection (MRP) method utilizes the group information in high dimensions to provide distinctive lower-dimensional projections that are then displayed using Radviz3D. Our methodology is extended to datasets with discrete and continuous features where a Gaussianized distributional transform is used in conjunction with copula models before applying MRP and visualizing the result using RadViz3D. A R package $radviz3d$ implementing our complete methodology is available.

stat.ML