SearcharxivSearch

arXiv subjects

D. Kugiumtzis

Publications and source records attributed to D. Kugiumtzis.

9 recordsLinked to original sources

Estimation of connectivity measures in gappy time series

A new method is proposed to compute connectivity measures on multivariate time series with gaps. Rather than removing or filling the gaps, the rows of the joint data matrix containing empty entries are removed and the calculations are done on the remainder matrix. The method, called measure adapted gap removal (MAGR), can be applied to any connectivity measure that uses a joint data matrix, such as cross correlation, cross mutual information and transfer entropy. MAGR is favorably compared using these three measures to a number of known gap-filling techniques, as well as the gap closure. The superiority of MAGR is illustrated on time series from synthetic systems and financial time series.

stat.ME

Reducing the Bias of Causality Measures

Measures of the direction and strength of the interdependence between two time series are evaluated and modified in order to reduce the bias in the estimation of the measures, so that they give zero values when there is no causal effect. For this, point shuffling is employed as used in the frame of surrogate data. This correction is not specific to a particular measure and it is implemented here on measures based on state space reconstruction and information measures. The performance of the causality measures and their modifications is evaluated on simulated uncoupled and coupled dynamical systems and for different settings of embedding dimension, time series length and noise level. The corrected measures, and particularly the suggested corrected transfer entropy, turn out to stabilize at the zero level in the absence of causal effect and detect correctly the direction of information flow when it is present. The measures are also evaluated on electroencephalograms (EEG) for the detection of the information flow in the brain of an epileptic patient. The performance of the measures on EEG is interpreted, in view of the results from the simulation study.

physics.data-an

Evaluation of mutual information estimators on nonlinear dynamic systems

Mutual information is a nonlinear measure used in time series analysis in order to measure the linear and non-linear correlations at any lag $τ$. The aim of this study is to evaluate some of the most commonly used mutual information estimators, i.e. estimators based on histograms (with fixed or adaptive bin size), $k$-nearest neighbors and kernels. We assess the accuracy of the estimators by Monte-Carlo simulations on time series from nonlinear dynamical systems of varying complexity. As the true mutual information is generally unknown, we investigate the existence and rate of consistency of the estimators (convergence to a stable value with the increase of time series length), and the degree of deviation among the estimators. The results show that the $k$-nearest neighbor estimator is the most stable and less affected by the method-specific parameter.

nlin.CD

State Space Reconstruction for Multivariate Time Series Prediction

In the nonlinear prediction of scalar time series, the common practice is to reconstruct the state space using time-delay embedding and apply a local model on neighborhoods of the reconstructed space. The method of false nearest neighbors is often used to estimate the embedding dimension. For prediction purposes, the optimal embedding dimension can also be estimated by some prediction error minimization criterion. We investigate the proper state space reconstruction for multivariate time series and modify the two abovementioned criteria to search for optimal embedding in the set of the variables and their delays. We pinpoint the problems that can arise in each case and compare the state space reconstructions (suggested by each of the two methods) on the predictive ability of the local model that uses each of them. Results obtained from Monte Carlo simulations on known chaotic maps revealed the non-uniqueness of optimum reconstruction in the multivariate case and showed that prediction criteria perform better when the task is prediction.

nlin.CD

Turning Point Prediction of Oscillating Time Series using Local Dynamic Regression Models

In the prediction of oscillating time series, the interest is in the turning points of successive oscillations rather than the samples themselves. For this purpose a scheme has been proposed; the state space reconstruction is limited to the turning points and the local (nearest neighbor) model is modified in order to predict the turning point magnitudes and times. This approach is extended here using a local dynamic regression model on both turning point magnitudes and times. Simulations on oscillating nonlinear systems show that the proposed approach gives better predictions of turning points than the standard local model applied to all the samples of the oscillating time series.

nlin.CD

Local prediction of turning points of oscillating time series

For oscillating time series, the prediction is often focused on the turning points. In order to predict the turning point magnitudes and times it is proposed to form the state space reconstruction only from the turning points and modify the local (nearest neighbor) model accordingly. The model on turning points gives optimal prediction at a lower dimensional state space than the optimal local model applied directly on the oscillating time series and is thus computationally more efficient. Monte Carlo simulations on different oscillating nonlinear systems showed that it gives better predictions of turning points and this is confirmed also for the time series of annual sunspots and total stress in a plastic deformation experiment.

nlin.CD

Statistical analysis of Gene and Intergenic DNA Sequences

Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate specifically on genes and intergenic regions. To estimate the level and type of correlation in these regions we apply various statistical methods inspired from nonlinear time series analysis, namely the probability distribution of tuplets, the Mutual Information and the Identical Neighbour Fit. The methods are suitably modified to work on symbolic sequences and they are first tested for validity on sequences obtained from well--known simple deterministic and stochastic models. Then they are applied to the DNA sequence of chromosome 1 of {\em arabidopsis thaliana}. The results suggest that correlations do exist in the DNA sequence but they are weak and that intergenic sequences tend to be more correlated than gene sequences. The use of statistical tests with surrogate data establish these findings in a rigorous statistical manner.

q-bio.GN

Statically Transformed Autoregressive Process and Surrogate Data Test for Nonlinearity

The key feature for the successful implementation of the surrogate data test for nonlinearity on a scalar time series is the generation of surrogate data that represent exactly the null hypothesis (statically transformed normal stochastic process), i.e. they possess the sample autocorrelation and amplitude distribution of the given data. A new conceptual approach and algorithm for the generation of surrogate data is proposed, called {\em statically transformed autoregressive process} (STAP). It identifies a normal autoregressive process and a monotonic static transform, so that the transformed realisations of this process fulfill exactly both conditions and do not suffer from bias in autocorrelation as the surrogate data generated by other algorithms. The appropriateness of STAP is demonstrated with simulated and real world data.

nlin.CD

Test your surrogate data before you test for nonlinearity

The schemes for the generation of surrogate data in order to test the null hypothesis of linear stochastic process undergoing nonlinear static transform are investigated as to their consistency in representing the null hypothesis. In particular, we pinpoint some important caveats of the prominent algorithm of amplitude adjusted Fourier transform surrogates (AAFT) and compare it to the iterated AAFT (IAAFT), which is more consistent in representing the null hypothesis. It turns out that in many applications with real data the inferences of nonlinearity after marginal rejection of the null hypothesis were premature and have to be re-investigated taken into account the inaccuracies in the AAFT algorithm, mainly concerning the mismatching of the linear correlations. In order to deal with such inaccuracies we propose the use of linear together with nonlinear polynomials as discriminating statistics. The application of this setup to some well-known real data sets cautions against the use of the AAFT algorithm.

physics.data-an