SearcharxivSearch

arXiv subjects

Alexander Kraskov

Publications and source records attributed to Alexander Kraskov.

9 recordsLinked to original sources

Independent components in spectroscopic analysis of complex mixtures

We applied two methods of "blind" spectral decomposition (MILCA and SNICA) to quantitative and qualitative analysis of UV absorption spectra of several non-trivial mixture types. Both methods use the concept of statistical independence and aim at the reconstruction of minimally dependent components from a linear mixture. We examined mixtures of major ecotoxicants (aromatic and polyaromatic hydrocarbons), amino acids and complex mixtures of vitamins in a veterinary drug. Both MICLA and SNICA were able to recover concentrations and individual spectra with minimal errors comparable with instrumental noise. In most cases their performance was similar to or better than that of other chemometric methods such as MCR-ALS, SIMPLISMA, RADICAL, JADE and FastICA. These results suggest that the ICA methods used in this study are suitable for real life applications. Data used in this paper along with simple matlab codes to reproduce paper figures can be found at http://www.klab.caltech.edu/~kraskov/MILCA/spectra

physics.chem-ph

MIC: Mutual Information based hierarchical Clustering

Clustering is a concept used in a huge variety of applications. We review a conceptually very simple algorithm for hierarchical clustering called in the following the {\it mutual information clustering} (MIC) algorithm. It uses mutual information (MI) as a similarity measure and exploits its grouping property: The MI between three objects X, Y, and Z is equal to the sum of the MI between X and Y, plus the MI between Z and the combined object (XY). We use MIC both in the Shannon (probabilistic) version of information theory, where the "objects" are probability distributions represented by random samples, and in the Kolmogorov (algorithmic) version, where the "objects" are symbol sequences. We apply our method to the construction of phylogenetic trees from mitochondrial DNA sequences and we reconstruct the fetal ECG from the output of independent components analysis (ICA) applied to the ECG of a pregnant woman.

q-bio.QM

Monte Carlo Algorithm for Least Dependent Non-Negative Mixture Decomposition

We propose a simulated annealing algorithm (called SNICA for "stochastic non-negative independent component analysis") for blind decomposition of linear mixtures of non-negative sources with non-negative coefficients. The de-mixing is based on a Metropolis type Monte Carlo search for least dependent components, with the mutual information between recovered components as a cost function and their non-negativity as a hard constraint. Elementary moves are shears in two-dimensional subspaces and rotations in three-dimensional subspaces. The algorithm is geared at decomposing signals whose probability densities peak at zero, the case typical in analytical spectroscopy and multivariate curve resolution. The decomposition performance on large samples of synthetic mixtures and experimental data is much better than that of traditional blind source separation methods based on principal component analysis (MILCA, FastICA, RADICAL) and chemometrics techniques (SIMPLISMA, ALS, BTEM) The source codes of SNICA, MILCA and the MI estimator are freely available online at http://www.fz-juelich.de/nic/cs/software

physics.chem-ph

Spectral Mixture Decomposition by Least Dependent Component Analysis

A recently proposed mutual information based algorithm for decomposing data into least dependent components (MILCA) is applied to spectral analysis, namely to blind recovery of concentrations and pure spectra from their linear mixtures. The algorithm is based on precise estimates of mutual information between measured spectra, which allows to assess and make use of actual statistical dependencies between them. We show that linear filtering performed by taking second derivatives effectively reduces the dependencies caused by overlapping spectral bands and, thereby, assists resolving pure spectra. In combination with second derivative preprocessing and alternating least squares postprocessing, MILCA shows decomposition performance comparable with or superior to specialized chemometrics algorithms. The results are illustrated on a number of simulated and experimental (infrared and Raman) mixture problems, including spectroscopy of complex biological materials. MILCA is available online at http://www.fz-juelich.de/nic/cs/software

physics.data-an

Least Dependent Component Analysis Based on Mutual Information

We propose to use precise estimators of mutual information (MI) to find least dependent components in a linearly mixed signal. On the one hand this seems to lead to better blind source separation than with any other presently available algorithm. On the other hand it has the advantage, compared to other implementations of `independent' component analysis (ICA) some of which are based on crude approximations for MI, that the numerical values of the MI can be used for: (i) estimating residual dependencies between the output components; (ii) estimating the reliability of the output, by comparing the pairwise MIs with those of re-mixed components; (iii) clustering the output according to the residual interdependencies. For the MI estimator we use a recently proposed k-nearest neighbor based algorithm. For time sequences we combine this with delay embedding, in order to take into account non-trivial time correlations. After several tests with artificial data, we apply the resulting MILCA (Mutual Information based Least dependent Component Analysis) algorithm to a real-world dataset, the ECG of a pregnant woman. The software implementation of the MILCA algorithm is freely available at http://www.fz-juelich.de/nic/cs/software

physics.comp-ph

Extracting Phases from Aperiodic Signals

We demonstrate by means of a simple example that the arbitrariness of defining a phase from an aperiodic signal is not just an academic problem, but is more serious and fundamental. Decomposition of the signal into components with positive phase velocities is proposed as an old solution to this new problem.

cond-mat.dis-nn

Hierarchical Clustering Based on Mutual Information

Motivation: Clustering is a frequently used concept in variety of bioinformatical applications. We present a new method for hierarchical clustering of data called mutual information clustering (MIC) algorithm. It uses mutual information (MI) as a similarity measure and exploits its grouping property: The MI between three objects X, Y, and Z is equal to the sum of the MI between X and Y, plus the MI between Z and the combined object (XY). Results: We use this both in the Shannon (probabilistic) version of information theory, where the "objects" are probability distributions represented by random samples, and in the Kolmogorov (algorithmic) version, where the "objects" are symbol sequences. We apply our method to the construction of mammal phylogenetic trees from mitochondrial DNA sequences and we reconstruct the fetal ECG from the output of independent components analysis (ICA) applied to the ECG of a pregnant woman. Availability: The programs for estimation of MI and for clustering (probabilistic version) are available at http://www.fz-juelich.de/nic/cs/software

q-bio.QM

Hierarchical Clustering Using Mutual Information

We present a method for hierarchical clustering of data called {\it mutual information clustering} (MIC) algorithm. It uses mutual information (MI) as a similarity measure and exploits its grouping property: The MI between three objects $X, Y,$ and $Z$ is equal to the sum of the MI between $X$ and $Y$, plus the MI between $Z$ and the combined object $(XY)$. We use this both in the Shannon (probabilistic) version of information theory and in the Kolmogorov (algorithmic) version. We apply our method to the construction of phylogenetic trees from mitochondrial DNA sequences and to the output of independent components analysis (ICA) as illustrated with the ECG of a pregnant woman.

q-bio.QM

Estimating Mutual Information

We present two classes of improved estimators for mutual information $M(X,Y)$, from samples of random points distributed according to some joint probability density $μ(x,y)$. In contrast to conventional estimators based on binnings, they are based on entropy estimates from $k$-nearest neighbour distances. This means that they are data efficient (with $k=1$ we resolve structures down to the smallest possible scales), adaptive (the resolution is higher where data are more numerous), and have minimal bias. Indeed, the bias of the underlying entropy estimates is mainly due to non-uniformity of the density at the smallest resolved scale, giving typically systematic errors which scale as functions of $k/N$ for $N$ points. Numerically, we find that both families become {\it exact} for independent distributions, i.e. the estimator $\hat M(X,Y)$ vanishes (up to statistical fluctuations) if $μ(x,y) = μ(x) μ(y)$. This holds for all tested marginal distributions and for all dimensions of $x$ and $y$. In addition, we give estimators for redundancies between more than 2 random variables. We compare our algorithms in detail with existing algorithms. Finally, we demonstrate the usefulness of our estimators for assessing the actual independence of components obtained from independent component analysis (ICA), for improving ICA, and for estimating the reliability of blind source separation.

cond-mat.stat-mech