SearcharxivSearch

arXiv subjects

Didier Fraix-Burnet

Publications and source records attributed to Didier Fraix-Burnet.

At least 19 recordsLinked to original sources

The Leinster-Cobbold diversity index as a criterion for sub-clustering

An automatic procedure to perform sub-clustering on large samples is presented. At each iteration, the most diverse cluster is sub-clustered, and the global diversity of the new classification is compared to the previous one. The process stops if no improvement is found. The key to our procedure is the use of a quantitative measure of diversity, called the Leinster-Cobbold index, that takes into account the similarity between clusters. While this procedure has been successfully applied on a large sample of spectra of galaxies, we illustrate its efficiency with two examples in this paper.

astro-ph.IM

Spectral similarities in galaxies through an unsupervised classification of spaxels

We present the first unsupervised classification of spaxels in hyperspectral images of individual galaxies. Classes identify regions by spectral similarity and thus take all the information into account that is contained in the data cubes (spatial and spectral).We used Gaussian mixture models in a latent discriminant subspace to find clusters of spaxels. The spectra were corrected for small-scale motions within the galaxy based on emission lines with an automatic algorithm. Our data consist of two MUSE/VLT data cubes of JKB 18 and NGC 1068 and one NIRSpec/JWST data cube of NGC 4151.Our classes identify many regions that are most often easily interpreted. Most of the 11 classes that we find for JKB 18 are identified as photoionised by stars. Some of them are known HII regions, but we mapped them as extended, with gradients of ionisation intensities. One compact structure has not been reported before, and according to diagnostic diagrams, it might be a planetary nebula or a denser HII region. For NGC 1068, our 16 classes are of active galactic nucleus-type (AGN) or star-forming regions. Their spatial distribution corresponds perfectly to well-known structures such as spiral arms and a ring with giant molecular clouds. A subclassification in the nuclear region reveals several structures and gradients in the AGN spectra. Our unsupervised classification of the MUSE data of NGC 1068 helps visualise the complex interaction of the AGN and the jet with the interstellar medium in a single map. The centre of NGC 4151 is very complex, but our classes can easily be related to ionisation cones, the jet, or H2 emission. We find a new elongated structure that is ionised by the AGN along the N-S axis perpendicular to the jet direction. It is rotated counterclockwise with respect to the axis of the H2 emission. Our work shows that the unsupervised classification of spaxels takes full advantage of the richness of the information in the data cubes by presenting the spectral and spatial information in a combined and synthetic way.

astro-ph.GA

Machine Learning and galaxy morphology: for what purpose?

Classification of galaxies is traditionally associated with their morphologies through visual inspection of images. The amount of data to come renders this task inhuman and Machine Learning (mainly Deep Learning) has been called to the rescue for more than a decade. However, the results look mitigate and there seems to be a shift away from the paradigm of the traditional morphological classification of galaxies. In this paper, I want to show that the algorithms indeed are very sensitive to the features present in images, features that do not necessarily correspond to the Hubble or de Vaucouleurs vision of a galaxy. However, this does not preclude to get the correct insights into the physics of galaxies. I have applied a state-of-the-art ''traditional'' Machine Learning clustering tool, called Fisher-EM, a latent discriminant subspace Gaussian Mixture Model algorithm, to 4458 galaxies carefully classified into 18 types by the EFIGI project. The optimum number of clusters given by the Integrated Complete Likelihood criterion is 47. The correspondence with the EFIGI classification is correct, but it appears that the Fisher-EM algorithm gives a great importance to the distribution of light which translates to characteristics such as the bulge to disk ratio, the inclination or the presence of foreground stars. The discrimination of some physical parameters (bulge-to-total luminosity ratio, $(B -- B)T$ , intrinsic diameter, presence of flocculence or dust, arm strength) is very comparable in the two classifications.

astro-ph.GA

Unsupervised classification of SDSS galaxy spectra

Defining templates of galaxy spectra is useful to quickly characterise new observations and organise databases from surveys. These templates are usually built from a pre-defined classification based on other criteria. Aims. We present an unsupervised classification of 702248 spectra of galaxies and quasars with redshifts smaller than 0.25 that were retrieved from the Sloan Digital Sky Survey (SDSS) database, release 7. The spectra were first corrected for redshift, then wavelet-filtered to reduce the noise, and finally binned to obtain about 1437 wavelengths per spectrum. The unsupervised clustering algorithm Fisher-EM, relying on a discriminative latent mixture model, was applied on these corrected spectra. The full set and several subsets of 100000 and 300000 spectra were analysed. The optimum number of classes given by a penalised likelihood criterion is 86 classes, of which the 37 most populated gather 99% of the sample. These classes are established from a subset of 302214 spectra. Using several cross-validation techniques we find that this classification agrees with the results obtained on the other subsets with an average misclassification error of about 15%. The large number of very small classes tends to increase this error rate. In this paper, we do an initial quick comparison of our classes with literature templates. This is the first time that an automatic, objective and robust unsupervised classification is established on such a large number of galaxy spectra. The mean spectra of the classes can be used as templates for a large majority of galaxies in our Universe.

astro-ph.GA

A Maximum Parsimony analysis of the effect of the environment on the evolution of galaxies

Context. Galaxy evolution and the effect of environment are most often studied using scaling relations or some regression analyses around some given property. These approaches however do not take into account the complexity of the physics of the galaxies and their diversification. Aims. We here investigate the effect of cluster environment on the evolution of galaxies through multivariate unsupervised classification and phylogenetic analyses applied to two relatively large samples from the WINGS survey, one of cluster members and one of field galaxies (2624 and 1476 objects respectively). Methods. These samples are the largest ones ever analysed with a phylogenetic approach in astrophysics. To be able to use the Maximum Parsimony (cladistics) method, we first performed a pre-clustering in 300 clusters with a hierarchical clustering technique, before applying it to these pre-clusters. All these computations used seven parameters: B-V, log(Re), nV , $μ$e , H$β$ , D4000 , log(M *). Results. We have obtained a tree for the combined samples and do not find different evolutionary paths for cluster and field galaxies. However, the cluster galaxies seem to have accelerated evolution in the sense they are statistically more diversified from a primitive common ancestor. The separate analyses show a hint for a slightly more regular evolution of the variables for the cluster galaxies, which may indicate they are more homogeneous as compared to field galaxies in the sense that the groups of the latter appear to have more specific properties. On the tree for the cluster galaxies, there is a separate branch which gathers rejunevated or stripped-off groups of galaxies. This branch is clearly visible on the colour-magnitude diagram, going back from the red sequence towards the blue one. On this diagram, the distribution and the evolutionary paths of galaxies are strikingly different for the two samples. Globally, we do not find any dominant variable able to explain either the groups or the tree structures. Rather, co-evolution appears everywhere, and could depend itself on environment or mass. Conclusions. This study is another demonstration that unsupervised machine learning is able to go beyond the simple scaling relations by taking into account several properties together. The phylogenetic approach is invaluable to trace the evolutionary scenarii and project them onto any biavariate diagram without any a priori modelling. Our WINGS galaxies are all at low redshift, and we now need to go to higher redshfits to find more primitve galaxies and complete the map of the evolutionary paths of present day galaxies.

astro-ph.GA

Unsupervised Classification of Galaxies. I. ICA feature selection

Subjective classification of galaxies can mislead us in the quest of the origin regarding formation and evolution of galaxies since this is necessarily limited to a few features. The human mind is not able to apprehend the complex correlations in a manyfold parameter space, and multivariate analyses are the best tools to understand the differences among various kinds of objects. In this series of papers, an objective classification of 362,923 galaxies from the Value Added Galaxy Catalogue (VAGC) is carried out with the help of two methods of multivariate analysis. First, Independent Component Analysis (ICA) is used to determine a set of derived independent components that are linear combinations of 47 observed features (viz. ionized lines, Lick indices, photometric and morphological properties, star formation rates etc.) of the galaxies. Subsequently, a K-means cluster analysis is applied on the nine independent components to obtain ten distinct and homogeneous groups. In this first paper, we describe the methods and the main results. It appears that the nine Independent Components represent a complete physical description of galaxies (velocity dispersion, ionisation, metallicity, surface brightness and structure). We find that our ten groups can be essentially placed into traditional and empirical classes (from colour-magnitude and emission-line diagnostic diagrams, early- vs late-types) despite the classical corresponding features (colour, line ratios and morphology) being not significantly correlated with the nine Independent Components. More detailed physical interpretation of the groups will be performed in subsequent papers.

astro-ph.CO

A phylogenetic approach to chemical tagging. Reassembling open cluster stars

Context. The chemical tagging technique is a promising approach to reconstruct the history of the Galaxy by only using stellar chemical abundances. Different studies have undertaken this analysis and they raised several challenges. Aims. Using a sample of open clusters stars, we wish to address two issues: minimize chemical abundance differences which origin is linked to the evolutionary stage of the stars and not their original composition; evaluate a phylogenetic approach to group stars based on their chemical composition. Methods. We derived differential chemical abundances for 207 stars (belonging to 34 open clusters) using the Sun as reference star (classical approach) and a dwarf plus a giant star from the open cluster M67 as reference (new approach). These abundances were then used to perform two phylogenetic analyses, cladistics (Maximum Parsimony) and Neighbour-Joining, together with a partitioning unsupervised classification analysis with k-means. The resulting groupings were finally confronted to the true open cluster memberships of the stars. Results. We successfully reconstruct most of the original open clusters when carefully selecting a subset of the abundances derived differentially with respect to M67. We find a set of eight chemical elements that yields the best result, and discuss the possible reasons for them to be good tracers of the history of the Galaxy. Conclusions. Our study shows that unraveling the history of the Galaxy by only using stellar chemical abundances is greatly improved provided that i) we perform a differential spectroscopic analysis with respect to an open cluster instead of the Sun, ii) select the chemical elements that are good tracers of the history of the Galaxy, and iii) use tools that are adapted to detect evolutionary tracks such as phylogenetic approaches.

astro-ph.SR

Phylogenetic Analyses of Quasars and Galaxies

Phylogenetic approaches have proven to be useful in astrophysics. We have recently published a Maximum Parsimony (or cladistics) analysis on two samples of 215 and 85 low-z quasars (z < 0.7) which offer a satisfactory coverage of the Eigenvector 1-derived main sequence. Cladistics is not only able to group sources radiating at higher Eddington ratios, to separate radio-quiet (RQ) and radio-loud (RL) quasars and properly distinguishes core-dominated and lobe-dominated quasars, but it suggests a black hole mass threshold for powerful radio emission as already proposed elsewhere. An interesting interpretation from this work is that the phylogeny of quasars may be represented by the ontogeny of their central black hole, i.e. the increase of the black hole mass. However these exciting results are based on a small sample of low-z quasars, so that the work must be extended. We are here faced with two difficulties. The first one is the current lack of a larger sample with similar observables. The second one is the prohibitive computation time to perform a cladistic analysis on more that about one thousand objects. We show in this paper an experimental strategy on about 1500 galaxies to get around this difficulty. Even if it not related to the quasar study, it is interesting by itself and opens new pathways to generalize the quasar findings.

astro-ph.GA

Phylogenetic Tools in Astrophysics

Multivariate clustering in astrophysics is a recent development justified by the bigger and bigger surveys of the sky. The phylogenetic approach is probably the most unexpected technique that has appeared for the unsupervised classification of galaxies, stellar populations or globular clusters. On one side, this is a somewhat natural way of classifying astrophysical entities which are all evolving objects. On the other side, several conceptual and practical difficulties arize, such as the hierarchical representation of the astrophysical diversity, the continuous nature of the parameters, and the adequation of the result to the usual practice for the physical interpretation. Most of these have now been solved through the studies of limited samples of stellar clusters and galaxies. Up to now, only the Maximum Parsimony (cladistics) has been used since it is the simplest and most general phylogenetic technique. Probabilistic and network approaches are obvious extensions that should be explored in the future.

astro-ph.IM

The phylogeny of quasars and the ontogeny of their central black holes

The connection between multifrequency quasar observational and physical parameters related to accretion processes is still open to debate. In the last 20 year, Eigenvector 1-based approaches developed since the early papers by Boroson and Green (1992) and Sulentic et al. (2000b) have been proven to be a remarkably powerful tool to investigate this issue, and have led to the definition of a quasar "main sequence". In this paper we perform a cladistic analysis on two samples of 215 and 85 low-z quasars (z 0.7) which were studied in several previous works and which offer a satisfactory coverage of the Eigenvector 1-derived main sequence. The data encompass accurate measurements of observational parameters which represent key aspects associated with the structural diversity of quasars. Cladistics is able to group sources radiating at higher Eddington ratios, as well as to separate radio-quiet (RQ) and radio-loud (RL) quasars. The analysis suggests a black hole mass threshold for powerful radio emission and also properly distinguishes core-dominated and lobe-dominated quasars, in accordance with the basic tenet of RL unification schemes. Considering that black hole mass provides a sort of "arrow of time" of nuclear activity, a phylogenetic interpretation becomes possible if cladistic trees are rooted on black hole mass: the ontogeny of black holes is represented by their monotonic increase in mass. More massive radio-quiet Population B sources at low-z become a more evolved counterpart of Population A i.e., wind dominated sources to which the "local" Narrow-Line Seyfert 1s belong.

astro-ph.GA

Concepts of Classification and Taxonomy. Phylogenetic Classification

Phylogenetic approaches to classification have been heavily developed in biology by bioinformaticians. But these techniques have applications in other fields, in particular in linguistics. Their main characteristics is to search for relationships between the objects or species in study, instead of grouping them by similarity. They are thus rather well suited for any kind of evolutionary objects. For nearly fifteen years, astrocladistics has explored the use of Maximum Parsimony (or cladistics) for astronomical objects like galaxies or globular clusters. In this lesson we will learn how it works. 1 Why phylogenetic tools in astrophysics? 1.1 History of classification The need for classifying living organisms is very ancient, and the first classification system can be dated back to the Greeks. The goal was very practical since it was intended to distinguish between eatable and toxic aliments, or kind and dangerous animals. Simple resemblance was used and has been used for centuries. Basically, until the XVIIIth century, every naturalist chose his own criterion to build a classification. At the end, hundreds of classifications were available, most often incompatible to each other. The criteria for this traditional way of classifying is the subjective appearance of the living organisms. During the XVIIIth a revolution occurred. Scientists like Adanson and Linn{é} devised new ways of classifying the objects and naming the classes. Adanson realised that all the observable traits should be used, giving birth to the mutivariate clustering and classification activity (Adanson, 1763). Linn{é} based his binomial nomenclature on neutral names unrelated whatsoever to any property of the classes. We can realise the success of these two ideas more than two centuries and a half later!

astro-ph.IM

Clustering with phylogenetic tools in astrophysics

Phylogenetic approaches are finding more and more applications outside the field of biology. Astrophysics is no exception since an overwhelming amount of multivariate data has appeared in the last twenty years or so. In particular, the diversification of galaxies throughout the evolution of the Universe quite naturally invokes phylogenetic approaches. We have demonstrated that Maximum Parsimony brings useful astrophysical results, and we now proceed toward the analyses of large datasets for galaxies. In this talk I present how we solve the major difficulties for this goal: the choice of the parameters, their discretization, and the analysis of a high number of objects with an unsupervised NP-hard classification technique like cladistics. 1. Introduction How do the galaxy form, and when? How did the galaxy evolve and transform themselves to create the diversity we observe? What are the progenitors to present-day galaxies? To answer these big questions, observations throughout the Universe and the physical modelisation are obvious tools. But between these, there is a key process, without which it would be impossible to extract some digestible information from the complexity of these systems. This is classification. One century ago, galaxies were discovered by Hubble. From images obtained in the visible range of wavelengths, he synthetised his observations through the usual process: classification. With only one parameter (the shape) that is qualitative and determined with the eye, he found four categories: ellipticals, spirals, barred spirals and irregulars. This is the famous Hubble classification. He later hypothetized relationships between these classes, building the Hubble Tuning Fork. The Hubble classification has been refined, notably by de Vaucouleurs, and is still used as the only global classification of galaxies. Even though the physical relationships proposed by Hubble are not retained any more, the Hubble Tuning Fork is nearly always used to represent the classification of the galaxy diversity under its new name the Hubble sequence (e.g. Delgado-Serrano, 2012). Its success is impressive and can be understood by its simplicity, even its beauty, and by the many correlations found between the morphology of galaxies and their other properties. And one must admit that there is no alternative up to now, even though both the Hubble classification and diagram have been recognised to be unsatisfactory. Among the most obvious flaws of this classification, one must mention its monovariate, qualitative, subjective and old-fashioned nature, as well as the difficulty to characterise the morphology of distant galaxies. The first two most significant multivariate studies were by Watanabe et al. (1985) and Whitmore (1984). Since the year 2005, the number of studies attempting to go beyond the Hubble classification has increased largely. Why, despite of this, the Hubble classification and its sequence are still alive and no alternative have yet emerged (Sandage, 2005)? My feeling is that the results of the multivariate analyses are not easily integrated into a one-century old practice of modeling the observations. In addition, extragalactic objects like galaxies, stellar clusters or stars do evolve. Astronomy now provides data on very distant objects, raising the question of the relationships between those and our present day nearby galaxies. Clearly, this is a phylogenetic problem. Astrocladistics 1 aims at exploring the use of phylogenetic tools in astrophysics (Fraix-Burnet et al., 2006a,b). We have proved that Maximum Parsimony (or cladistics) can be applied in astrophysics and provides a new exploration tool of the data (Fraix-Burnet et al., 2009, 2012, Cardone \& Fraix-Burnet, 2013). As far as the classification of galaxies is concerned, a larger number of objects must now be analysed. In this paper, I

astro-ph.IM

Multivariate Approaches to Classification in Extragalactic Astronomy

Clustering objects into synthetic groups is a natural activity of any science. Astrophysics is not an exception and is now facing a deluge of data. For galaxies, the one-century old Hubble classification and the Hubble tuning fork are still largely in use, together with numerous mono-or bivariate classifications most often made by eye. However, a classification must be driven by the data, and sophisticated multivariate statistical tools are used more and more often. In this paper we review these different approaches in order to situate them in the general context of unsupervised and supervised learning. We insist on the astrophysical outcomes of these studies to show that multivariate analyses provide an obvious path toward a renewal of our classification of galaxies and are invaluable tools to investigate the physics and evolution of galaxies.

astro-ph.GA

Stellar populations in $ω$ Centauri: a multivariate analysis

We have performed multivariate statistical analyses of photometric and chemical abundance parameters of three large samples of stars in the globular cluster $ω$ Centauri. The statistical analysis of a sample of 735 stars based on seven chemical abundances with the method of Maximum Parsimony (cladistics) yields the most promising results: seven groups are found, distributed along three branches with distinct chemical, spatial and kinematical properties. A progressive chemical evolution can be traced from one group to the next, but also within groups, suggestive of an inhomogeneous chemical enrichment of the initial interstellar matter. The adjustment of stellar evolution models shows that the groups with metallicities [Fe/H]\textgreater{}-1.5 are Helium-enriched, thus presumably of second generation. The spatial concentration of the groups increases with chemical evolution, except for two groups, which stand out in their other properties as well. The amplitude of rotation decreases with chemical evolution, except for two of the three metal-rich groups, which rotate fastest, as predicted by recent hydrodynamical simulations. The properties of the groups are interpreted in terms of star formation in gas clouds of different origins. In conclusion, our multivariate analysis has shown that metallicity alone cannot segregate the different populations of $ω$ Centauri.

astro-ph.GA

Clustering large number of extragalactic spectra of galaxies and quasars through canopies

Cluster analysis is the distribution of objects into different groups or more precisely the partitioning of a data set into subsets (clusters) so that the data in subsets share some common trait according to some distance measure. Unlike classi cation, in clustering one has to rst decide the optimum number of clusters and then assign the objects into different clusters. Solution of such problems for a large number of high dimensional data points is quite complicated and most of the existing algorithms will not perform properly. In the present work a new clustering technique applicable to large data set has been used to cluster the spectra of 702248 galaxies and quasars having 1540 points in wavelength range imposed by the instrument. The proposed technique has successfully discovered ve clusters from this 702248X1540 data matrix.

astro-ph.CO

A six-parameter space to describe galaxy diversification

Galaxy diversification proceeds by transforming events like accretion, interaction or mergers. These explain the formation and evolution of galaxies that can now be described with many observables. Multivariate analyses are the obvious tools to tackle the datasets and understand the differences between different kinds of objects. However, depending on the method used, redundancies, incompatibilities or subjective choices of the parameters can void the usefulness of such analyses. The behaviour of the available parameters should be analysed before an objective reduction of dimensionality and subsequent clustering analyses can be undertaken, especially in an evolutionary context. We study a sample of 424 early-type galaxies described by 25 parameters, ten of which are Lick indices, to identify the most structuring parameters and determine an evolutionary classification of these objects. Four independent statistical methods are used to investigate the discriminant properties of the observables and the partitioning of the 424 galaxies: Principal Component Analysis, K-means cluster analysis, Minimum Contradiction Analysis and Cladistics. (abridged)

astro-ph.CO

Multivariate Evolutionary Analyses in Astrophysics

The large amount of data on galaxies, up to higher and higher redshifts, asks for sophisticated statistical approaches to build adequate classifications. Multivariate cluster analyses, that compare objects for their global similarities, are still confidential in astrophysics, probably because their results are somewhat difficult to interpret. We believe that the missing key is the unavoidable characteristics in our Universe: evolution. Our approach, known as Astrocladistics, is based on the evolutionary nature of both galaxies and their properties. It gathers objects according to their "histories" and establishes an evolutionary scenario among groups of objects. In this presentation, I show two recent results on globular clusters and earlytype galaxies to illustrate how the evolutionary concepts of Astrocladistics can also be useful for multivariate analyses such as K-means Cluster Analysis.

astro-ph.CO

The Fundamental Plane of Early-Type Galaxies as a Confounding Correlation

Early-type galaxies are characterized by many scaling relations. One of them, the so-called fundamental plane is a relatively tight correlation between three variables, and has resisted a clear physical understanding despite many years of intensive research. Here, we show that the correlation between the three variables of the fundamental plane can be the artifact of the effect of another parameter influencing all, so that the fundamental plane may be understood as a confounding correlation. Indeed, the complexity of the physics of galaxies and of their evolution suggests that the main confounding parameter must be related to the level of diversification reached by the galaxies. Consequently, many scaling relations for galaxies are probably evolutionary correlations.

astro-ph.CO