SearcharxivSearch

arXiv subjects

Soumita Modak

Publications and source records attributed to Soumita Modak.

14 recordsLinked to original sources

A new completely parameter-free clustering algorithm for unsupervised classification of BATSE gamma-ray bursts

Cluster analysis is a widely applied machine learning technique to understand the existing patterns in the population of gamma-ray bursts (GRBs), in order to explore their physical sources. In the present scenario, the number of clusters corresponding to differentiable groups is still under conflict, in spite of numerous attempts with the state-of-the-art clustering procedures. This crucial unknown parameter needs to be evaluated, either directly or indirectly in terms of other tuning parameters, to produce the clusters in GRBs through implementation of an appropriate clustering algorithm. While most of the applied algorithms reached two physically explained groups of merger and collapsar predominated by the short and long bursts respectively, other statistical approaches violated this binary partition. However, physical establishment of any additional cluster(s) is not yet confirmed. Therefore, we propose a new algorithm, from a different stream of clustering referred to as `completely parameter-free', which carries out the classification of GRBs in a manner that has not been tried so far. It indicates two main groups, of short and long duration bursts from the BATSE sample, compatible with the merger-collapsar theory.

astro-ph.HE

Evaluation of the number of clusters in a data set using $p$-values from Multiple Tests of Hypotheses

This paper proposes a novel, nonparametric, interpoint distance-based measure to investigate whether there exist any groups in a set of given data, and if so then, how many groups are prevailing in total. It is a cluster accuracy index useful for arbitrary-dimensional data set, in association with any clustering algorithm having the number of groups specified as a priori. We perform univariate, nonparametric, multiple statistical tests of hypotheses, where as many dependent tests as the sample size are carried out using the interpoint distances. They possess $p$-values to be combined to reach a decision, which is taken in a step-wise process for a possible number of clusters. It reduces the unnecessary computations compared with the other accuracy measures from the literature. Data study establishes the proposed index's efficiency and superiority.

stat.ME

Confirmation of Binary Clustering in Gamma-Ray Bursts through an Integrated $p$-value from Multiple Nonparametric Tests of Hypotheses

The paper applies a new, nonparametric, interpoint distance-based measure to confirm the inherent groups prevailing in the brightest source of light in the universe: gamma-ray bursts. Our effective metric, in association with clustering methods like Gaussian-mixture model-based and $K$-means algorithms, resolves the conflict regarding the possibility about existence of more than binary clusters in the gamma-ray burst population. Here we carry out multiple nonparametric statistical tests of hypotheses, as many as the number of bursts available from the `BATSE' catalog. An integrated $p$-value achieved from the aforesaid dependent tests solves our concern confirming two groups of short and long bursts.

astro-ph.HE

Quality check of a sample partition using multinomial distribution

In this paper, we advocate a novel measure for the purpose of checking the quality of a cluster partition for a sample into several distinct classes, and thus, determine the unknown value for the true number of clusters prevailing the provided set of data. Our objective leads us to the development of an approach through applying the multinomial distribution to the distances of data members, clustered in a group, from their respective cluster representatives. This procedure is carried out independently for each of the clusters, and the concerned statistics are combined together to design our targeted measure. Individual clusters separately possess the category-wise probabilities which correspond to different positions of its members in the cluster with respect to a typical member, in the form of cluster-centroid, medoid or mode, referred to as the corresponding cluster representative. Our method is robust in the sense that it is distribution-free, since this is devised irrespective of the parent distribution of the underlying sample. It fulfills one of the rare coveted qualities, present in the existing cluster accuracy measures, of having the capability to investigate whether the assigned sample owns any inherent clusters other than a single group of all members or not. Our measure's simple concept, easy algorithm, fast runtime, good performance, and wide usefulness, demonstrated through extensive simulation and diverse case-studies, make it appealing.

stat.AP

A new interpoint distance-based clustering algorithm using kernel density estimation

A novel nonparametric clustering algorithm is proposed using the interpoint distances between the members of the data to reveal the inherent clustering structure existing in the given set of data, where we apply the classical nonparametric univariate kernel density estimation method to the interpoint distances to estimate the density around a data member. Our clustering algorithm is simple in its formation and easy to apply resulting in well-defined clusters. The algorithm starts with objective selection of the initial cluster representative and always converges independently of this choice. The method finds the number of clusters itself and can be used irrespective of the nature of underlying data by using an appropriate interpoint distance measure. The cluster analysis can be carried out in any dimensional space with viability to high-dimensional use. The distributions of the data or their interpoint distances are not required to be known due to the design of our procedure, except the assumption that the interpoint distances possess a density function. Data study shows its effectiveness and superiority over the widely used clustering algorithms.

stat.ME

A new nonparametric interpoint distance-based measure for assessment of clustering

A new interpoint distance-based measure is proposed to identify the optimal number of clusters present in a data set. Designed in nonparametric approach, it is independent of the distribution of given data. Interpoint distances between the data members make our cluster validity index applicable to univariate and multivariate data measured on arbitrary scales, or having observations in any dimensional space where the number of study variables can be even larger than the sample size. Our proposed criterion is compatible with any clustering algorithm, and can be used to determine the unknown number of clusters or to assess the quality of the resulting clusters for a data set. Demonstration through synthetic and real-life data establishes its superiority over the well-known clustering accuracy measures of the literature.

cs.LG

Clustering of eclipsing binary light curves through functional principal component analysis

In this paper, we revisit the problem of clustering 1318 new variable stars found in the Milky way. Our recent work distinguishes these stars based on their light curves which are univariate series of brightness from the stars observed at discrete time points. This work proposes a new approach to look at these discrete series as continuous curves over time by transforming them into functional data. Then, functional principal component analysis is performed using these functional light curves. Clustering based on the significant functional principal components reveals two distinct groups of eclipsing binaries with consistency and superiority compared to our previous results. This method is established as a new powerful light curve-based classifier, where implementation of a simple clustering algorithm is effective enough to uncover the true clusters based merely on the first few relevant functional principal components. Simultaneously we discard the noise from the data study involving the higher order functional principal components. Thus the suggested method is very useful for clustering big light curve data sets which is also verified by our simulation study.

stat.AP

A new measure for assessment of clustering based on kernel density estimation

A new clustering accuracy measure is proposed to determine the unknown number of clusters and to assess the quality of clustering of a data set given in any dimensional space. Our validity index applies the classical nonparametric univariate kernel density estimation method to the interpoint distances computed between the members of data. Being based on interpoint distances, it is free of the curse of dimensionality and therefore efficiently computable for high-dimensional situations where the number of study variables can be larger than the sample size. The proposed measure is compatible with any clustering algorithm and with every kind of data set where the interpoint distance measure can be defined to have a density function. Simulation study proves its superiority over widely used cluster validity indices like the average silhouette width and the Dunn index, whereas its applicability is shown with respect to a high-dimensional Biostatistical study of Alon data set and a large Astrostatistical application of time series with light curves of new variable stars.

stat.ME

Distinction of groups of gamma-ray bursts in the BATSE catalog through fuzzy clustering

In search for the possible astrophysical sources behind origination of the diverse gamma-ray bursts, cluster analyses are performed to find homogeneous groups, which discover an intermediate group other than the conventional short and long bursts. However, very recently, few studies indicate a possibility of the existence of more than three (namely five) groups. Therefore, in this paper, fuzzy clustering is conducted on the gamma-ray bursts from the final 'Burst and Transient Source Experiment' catalog to cross-check the significance of these new groups. Meticulous study on individual bursts based on their memberships in the fuzzy clusters confirms the previously well-known three groups against the newly found five.

stat.AP

Unsupervised classification of eclipsing binary light curves through k-medoids clustering

This paper proposes k-medoids clustering method to reveal the distinct groups of 1,318 variable stars in the Galaxy based on their light curves, where each light curve represents the graph of brightness of the star against time. To overcome the deficiencies of subjective traditional classification, we separate the stars more scientifically according to their geometrical configuration and show that our approach outperforms the existing classification schemes in astronomy. It results in two optimum groups of eclipsing binaries corresponding to bright, massive systems and fainter, less massive systems.

astro-ph.SR

Bivariate density estimation using normal-gamma kernel with application to astronomy

We consider the problem of estimation of a bivariate density function with support $\Re\times[0,\infty)$, where a classical bivariate kernel estimator causes boundary bias due to the non-negative variable. To overcome this problem, we propose four kernel density estimators and compare their performances in terms of the mean integrated squared error. Simulation study shows that the estimator based on the proposed normal-gamma kernel performs best. Two astronomical data sets are used to demonstrate the applicability of this estimator.

stat.AP

A new nonparametric test for two sample multivariate location problem with application to astronomy

This paper provides a nonparametric test for the identity of two multivariate continuous distribution functions (d.f.'s) when they differ in locations. The test uses Wilcoxon rank-sum statistics on distances between observations for each of the components and is unaffected by outliers. It is numerically compared with two existing procedures in terms of power. The simulation study shows that its power is strictly increasing in the sample sizes and/or in the number of components. The applicability of this test is demonstrated by use of two astronomical data sets on early-type galaxies.

stat.AP

Clustering of Gamma-Ray bursts through kernel principal component analysis

We consider the problem related to clustering of gamma-ray bursts (from "BATSE" catalogue) through kernel principal component analysis in which our proposed kernel outperforms results of other competent kernels in terms of clustering accuracy and we obtain three physically interpretable groups of gamma-ray bursts. The effectivity of the suggested kernel in combination with kernel principal component analysis in revealing natural clusters in noisy and nonlinear data while reducing the dimension of the data is also explored in two simulated data sets.

stat.AP

Two phase formation of massive elliptical galaxies : study through cross-correlation including spatial effect

Formation mechanism of present day population of elliptical galaxies have been revisited in the context of hierarchical cosmological models accompanied by accretion and minor mergers through cross correlation function including spatial effect. The present work investigates the formation and evolution of several components of nearby massive early type galaxies (ETGs) through cross-correlation in the spatial coordinates, right ascension and declination (RA, DEC) and mass-size parameter space with high redshift $(0.5\leq z\leq2.7)$ ETGs. It is found that innermost components of nearby ETGs are highly correlated with ETGs in the redshift range $(2\leq z\leq2.7)$ known as 'red nuggets'. The intermediate and outermost parts have moderate correlations with ETGs in the redshift range $(0.5\leq z\leq0.75)$. The quantitative measures are highly consistent with the two phase formation scenario of massive nearby early type galaxies as suggested by various authors and resolves the conflict raised in a previous work suggesting other possibilities for the formation of outermost part of nearby massive ETGs. The improvement is expected to be due to inclusion of spatial effects in addition to other linear parameters.

astro-ph.GA