SearcharxivSearch

arXiv subjects

Radhakrishnan Nagarajan

Publications and source records attributed to Radhakrishnan Nagarajan.

At least 19 recordsLinked to original sources

Deciphering Dynamical Nonlinearities in Short Time Series Using Recurrent Neural Networks

Surrogate testing techniques have been used widely to investigate the presence of dynamical nonlinearities, an essential ingredient of deterministic chaotic processes. Traditional surrogate testing subscribes to statistical hypothesis testing and investigates potential differences in discriminant statistics between the given empirical sample and its surrogate counterparts. The choice and estimation of the discriminant statistics can be challenging across short time series. Also, conclusion based on a single empirical sample is an inherent limitation. The present study proposes a recurrent neural network classification framework that uses the raw time series obviating the need for discriminant statistic while accommodating multiple time series realizations for enhanced generalizability of the findings. The results are demonstrated on short time series with lengths (L = 32, 64, 128) from continuous and discrete dynamical systems in chaotic regimes, nonlinear transform of linearly correlated noise and experimental data. Accuracy of the classifier is shown to be markedly higher than >> 50% for the processes in chaotic regimes whereas those of nonlinearly correlated noise were around ~50% similar to that of random guess from a one-sample binomial test. These results are promising and elucidate the usefulness of the proposed framework in identifying potential dynamical nonlinearities from short experimental time series.

eess.SP

Network Abstractions of Prescription Patterns in a Medicaid Population

Understanding prescription patterns have relied largely on aggregate statistical measures. Evidence of doctor-shopping, inappropriate prescribing, drug diversion and patient seeking prescription drugs across multiple prescribers demand understanding the concerted working of prescribers and prescriber communities as opposed to treating them as independent entities. We model potential associations between prescribers as prescriber-prescriber network (PPN) and subsequently investigate its properties across Schedule II, III, IV drugs in a single month in a Medicaid population. Community structure detection algorithms and geo-spatial layouts revealed characteristic patterns in PPN markedly different from their random graph surrogate counterparts rejecting them as potential generative mechanism. Outlier detection with recommended thresholds also revealed a subset of prescriber specialties to be constitutively flagged across Schedule II, III, IV drugs. Presence of prescriber communities may assist in targeted monitoring and their deviation from random graphs may serve as a metric in assessing PPN evolution temporally and pre-/post- interventions.

stat.AP

On Identifying Significant Edges in Graphical Models of Molecular Networks

Objective: Modelling the associations from high-throughput experimental molecular data has provided unprecedented insights into biological pathways and signalling mechanisms. Graphical models and networks have especially proven to be useful abstractions in this regard. Ad-hoc thresholds are often used in conjunction with structure learning algorithms to determine significant associations. The present study overcomes this limitation by proposing a statistically-motivated approach for identifying significant associations in a network. Methods and Materials: A new method that identifies significant associations in graphical models by estimating the threshold minimising the $L_{\mathrm{1}}$ norm between the cumulative distribution function (CDF) of the observed edge confidences and those of its asymptotic counterpart is proposed. The effectiveness of the proposed method is demonstrated on popular synthetic data sets as well as publicly available experimental molecular data corresponding to gene and protein expression profiles. Results: The improved performance of the proposed approach is demonstrated across the synthetic data sets using sensitivity, specificity and accuracy as performance metrics. The results are also demonstrated across varying sample sizes and three different structure learning algorithms with widely varying assumptions. In all cases, the proposed approach has specificity and accuracy close to 1, while sensitivity increases linearly in the logarithm of the sample size. The estimated threshold systematically outperforms common ad-hoc ones in terms of sensitivity while maintaining comparable levels of specificity and accuracy. Networks from experimental data sets are reconstructed accurately with respect to the results from the original papers.

stat.ML

Comment of Global dynamics of biological systems

In a recent study, (Grigorov, 2006) analyzed temporal gene expression profiles (Arbeitman et al., 2002) generated in a Drosophila experiment using SSA in conjunction with Monte-Carlo SSA. The author (Grigorov, 2006) makes three important claims in his article, namely: Claim1: A new method based on the theory of nonlinear time series analysis is used to capture the global dynamics of the fruit-fly cycle temporal gene expression profiles. Claim 2: Flattening of a significant part of the eigen-spectrum confirms the hypothesis about an underly-ing high-dimensional chaotic generating process. Claim 3: Monte-Carlo SSA can be used to establish whether a given time series is distinguishable from any well-defined process including deterministic chaos. In this report we present fundamental concerns with respect to the above claims (Grigorov, 2006) in a systematic manner with simple examples. The discussion provided especially discourages the choice of SSA for inferring nonlinear dynamical structure form time series obtained in any biological paradigm.

q-bio.GN

Information retrieval from a phoneme time series database

Developing fast and efficient algorithms for retrieval of objects to a given user query is an area of active research. The present study investigates retrieval of time series objects from a phoneme database to a given user pattern or query. The proposed method maps the one-dimensional time series retrieval into a sequence retrieval problem by partitioning the multi-dimensional phase-space using k-means clustering. The problem of whole sequence as well as subsequence matching is considered. Robustness of the proposed technique is investigated on phoneme time series corrupted with additive white Gaussian noise. The shortcoming of classical power-spectral techniques for time series retrieval is also discussed.

q-bio.QM

Power-law Signatures and Patchiness in Genechip Oligonucleotide Microarrays

. Genechip oligonucleotide microarrays have been used widely for transcriptional profiling of a large number of genes in a given paradigm. Gene expression estimation precedes biological inference and is given as a complex combination of atomic entities on the array called probes. These probe intensities are further classified into perfect-match (PM) and mis-match (MM) probes. While former is a measure of specific binding, the lat-ter is a measure of non-specific binding. The behavior of the MM probes has especially proven to be elusive. The present study investigates qualita-tive similarities in the distributional signatures and local correlation struc-tures/patchiness between the PM and MM probe intensities. These qualita-tive similarities are established on publicly available microarrays generated across laboratories investigating the same paradigm. Persistence of these similarities across raw as well as background subtracted probe intensities is also investigated. The results presented raise fundamental concerns in inter-preting Genechip oligonucleotide microarray data.

q-bio.GN

Interpreting non-random signatures in biomedical signals with Lempel-Ziv complexity

Lempel-Ziv complexity (LZ) [1] and its variants have been used widely to identify non-random patterns in biomedical signals obtained across distinct physiological states. Non-random signatures of the complexity measure can occur under nonlinear deterministic as well as non-deterministic settings. Surrogate data testing have also been encouraged in the past in conjunction with complexity estimates to make a finer distinction between various classes of processes. In this brief letter, we make two important observations (1) Non-Gaussian noise at the dynamical level can elude existing surrogate algorithms namely: Phase-randomized surrogates (FT) amplitude-adjusted Fourier transform (AAFT) and iterated amplitude adjusted Fourier transform (IAAFT). Thus any inference nonlinear determinism as an explanation for the non-randomness is incomplete (2) Decrease in complexity can be observed even across two linear processes with identical auto-correlation functions. The results are illustrated with a second-order auto-regressive process with Gaussian and non-Gaussian innovations. AR (2) processes have been used widely to model several physiological phenomena, hence their choice. The results presented encourage cautious interpretation of non-random signatures in experimental signals using complexity measures.

nlin.CD

Delay estimation in a two-node acyclic network

Linear measures such as cross-correlation have been used successfully to determine time delays from the given processes. Such an analysis often precedes identifying possible causal relationships between the observed processes. The present study investigates the impact of a positively correlated driver whose correlation function decreases monotonically with lag on the delay estimation in a two-node acyclic network with one and two-delays. It is shown that cross-correlation analysis of the given processes can result in spurious identification of multiple delays between the driver and the dependent processes. Subsequently, delay estimation of increment process as opposed to the original process under certain implicit constraints is explored. Short-range and long-range correlated driver processes along with those of their coarse-grained counterparts are considered.

q-bio.QM

Surrogate testing of volatility series from long-range correlated noise

Detrended fluctuation analysis (DFA) [1] of the volatility series has been found to be useful in dentifying possible nonlinear/multifractal dynamics in the empirical sample [2-4]. Long-range volatile correlation can be an outcome of static as well as dynamical nonlinearity. In order to argue in favor of dynamical nonlinearity, surrogate testing is used in conjunction with volatility analysis [2-4]. In this brief communication, surrogate testing of volatility series from long-range correlated noise and their static, invertible nonlinear transforms is investigated. Long-range correlated monofractal noise is generated using FARIMA (0, d, 0) with Gaussian and non-Gaussian innovations. We show significant deviation in the scaling behavior between the empirical sample and the surrogate counterpart at large time-scales in the case of FARIMA (0, d, 0) with non-Gaussian innovations whereas no such discrepancy was observed in the case of Gaussian innovations. The results encourage cautious interpretation of surrogate testing in the presence of non-Gaussian innovations.

physics.data-an

Qualitative Assessment of Gene Expression in Affymetrix Genechip Arrays

Affymetrix Genechip microarrays are used widely to determine the simultaneous expression of genes in a given biological paradigm. Probes on the Genechip array are atomic entities which by definition are randomly distributed across the array and in turn govern the gene expression. In the present study, we make several interesting observations. We show that there is considerable correlation between the probe intensities across the array which defy the independence assumption. While the mechanism behind such correlations is unclear, we show that scaling behavior and the profiles of perfect match (PM) as well as mismatch (MM) probes are similar and immune to background subtraction. We believe that the observed correlations are possibly an outcome of inherent non-stationarities or patchiness in the array devoid of biological significance. This is demonstrated by inspecting their scaling behavior and profiles of the PM and MM probe intensities obtained from publicly available Genechip arrays from three eukaryotic genomes, namely: Drosophila Melanogaster, Homo Sapiens and Mus musculus across distinct biological paradigms and across laboratories, with and without background subtraction. The fluctuation functions were estimated using detrended fluctuation analysis (DFA) with fourth order polynomial detrending. The results presented in this study provide new insights into correlation signatures of PM and MM probe intensities and suggests the choice of DFA as a tool for qualitative assessment of Affymetrix Genechip microarrays prior to their analysis. A more detailed investigation is necessary in order to understand the source of these correlations.

q-bio.GN

A Qualitative Description of Boundary Layer Wind Speed Records

The complexity of the atmosphere endows it with the property of turbulence by virtue of which, wind speed variations in the atmospheric boundary layer (ABL) exhibit highly irregular fluctuations that persist over a wide range of temporal and spatial scales. Despite the large and significant body of work on microscale turbulence, understanding the statistics of atmospheric wind speed variations has proved to be elusive and challenging. Knowledge about the nature of wind speed at ABL has far reaching impact on several fields of research such as meteorology, hydrology, agriculture, pollutant dispersion, and more importantly wind energy generation. In the present study, temporal wind speed records from twenty eight stations distributed through out the state of North Dakota (ND, USA), ($\sim$ 70,000 square-miles) and spanning a period of nearly eight years are analyzed. We show that these records exhibit a characteristic broad multifractal spectrum irrespective of the geographical location and topography. The rapid progression of air masses with distinct qualitative characteristics originating from Polar regions, Gulf of Mexico and Northern Pacific account for irregular changes in the local weather system in ND. We hypothesize that one of the primary reasons for the observed multifractal structure could be the irregular recurrence and confluence of these three air masses.

physics.ao-ph

Synchronization in Electrically Coupled Neural Networks

In this report, we investigate the synchronization of temporal activity in an electrically coupled neural network model. The electrical coupling is established by homotypic static gap-junctions (Connexin 43). Two distinct network topologies, namely: {\em sparse random network, (SRN)} and {\em fully connected network, (FCN)} are used to establish the connectivity. The strength of connectivity in the FCN is governed by the {\em mean gap junctional conductance} ($μ$). In the case of the SRN, the overall strength of connectivity is governed by the {\em density of connections} ($δ$) and the connection strength between two neurons ($S_0$). The synchronization of the network with increasing gap junctional strength and varying population sizes is investigated. It was observed that the network {\em abruptly} makes a transition from a weakly synchronized to a well synchronized regime when ($δ$) or ($μ$) exceeds a critical value. It was also observed that the ($δ$, $μ$) values used to achieve synchronization decreases with increasing network size.

q-bio.NC

Surrogate testing of linear feedback processes with non-Gaussian innovations

Surrogate testing is used widely to determine the nature of the process generating the given empirical sample. In the present study, the usefulness of phase-randomized surrogates, amplitude adjusted Fourier transform (AAFT) and iterated amplitude adjusted Fourier transform (IAAFT) surrogates on statistical inference of linearly correlated noise with non-Gaussian innovations and their static, invertible nonlinear transforms from their empirical samples is discussed. Existing surrogate testing procedures which retain the auto-correlation function in the surrogates may not be appropriate in the presence of non-Gaussian innovations.

cond-mat.stat-mech

Is non-Gaussianity sufficient to produce long-range volatile correlations?

Scaling analysis of the magnitude series (volatile series) has been proposed recently to identify possible nonlinear/multifractal signatures in the given data [1-3]. In this letter, correlations of volatile series generated from stationary first-order linear feedback process with Gaussian and non-Gaussian innovations are investigated. While volatile correlations corresponding to Gaussian innovations exhibited uncorrelated behavior across all time scales, those of non-Gaussian innovations showed significant deviation from uncorrelated behavior even at large time scales. The results presented raise the intriguing question whether non-Gaussian innovations can be sufficient to realize long-range volatile correlations.

cond-mat.stat-mech

Appropriateness of correlated first order auto-regressive processes for modeling daily temperature records

The present study investigates linear and volatile (nonlinear) correlations of first-order autoregressive process with uncorrelated AR (1) and long-range correlated CAR (1) Gaussian innovations as a function of the process parameter ($θ$). In the light of recent findings \cite{jano}, we discuss the choice of CAR (1) in modeling daily temperature records. We demonstrate that while CAR (1) is able to capture linear correlations it is unable to capture nonlinear (volatile) correlations in daily temperature records.

physics.ao-ph

Correlation Statistics for cDNA Microarray Image Analysis

In this report, correlation of the pixels comprising a microarray spot is investigated. Subsequently, correlation statistics namely: Pearson correlation and Spearman rank correlation are used to segment the foreground and background intensity of microarray spots. The performance of correlation-based segmentation is compared to clustering-based (PAM, k-means) and seeded-region growing techniques (SPOT). It is shown that correlation-based segmentation is useful in flagging poorly hybridized spots, thus minimizes false-positives. The present study also raises the intriguing question of whether a change in correlation can be an indicator of differential gene expression.

q-bio.GN

Local Analysis of Dissipative Dynamical Systems

Linear transformation techniques such as singular value decomposition (SVD) have been used widely to gain insight into the qualitative dynamics of data generated by dynamical systems. There have been several reports in the past that had pointed out the susceptibility of linear transformation approaches in the presence of nonlinear correlations. In this tutorial review, local dispersion along with the surrogate testing is proposed to discriminate nonlinear correlations arising in deterministic and non-deterministic settings.

nlin.CD

Impact of Tandem Repeats on the Scaling of Nucleotide Sequences

Techniques such as detrended fluctuation analysis (DFA) and its extensions have been widely used to determine the nature of scaling in nucleotide sequences. In this brief communication we show that tandem repeats which are ubiquitous in nucleotide sequences can prevent reliable estimation of possible long-range correlations. Therefore, it is important to investigate the presence of tandem repeats prior to scaling exponent estimation.

q-bio.GN