SearcharxivSearch

arXiv subjects

Sarbojit Roy

Publications and source records attributed to Sarbojit Roy.

6 recordsLinked to original sources

KenCoh: A Ranked-Based Canonical Coherence

This work is inspired by the problem of characterizing a dependence measure between two cortical regions of the brain where each region contains multiple signal recordings from several neurons or channels (e.g., inhibitory and excitatory neurons). The goal is to identify differences in the structure of brain functional connectivity between known brain states. An exploratory tool for studying the dependence between two random vectors is via canonical correlation analysis. However, these are limited to only capturing linear associations and are sensitive to outlier observations. Mitigating these limitations is crucial because brain functional connectivity is likely to be more complex than linear, and brain signals may exhibit heavy-tailed properties. To overcome these limitations, we develop a robust method, Kendall's tau-based canonical coherence (KenCoh), to learn connectivity structure among neuronal signals filtered at given frequency bands. Our simulation study demonstrates that KenCoh is competitive with the moment-based estimator and outperforms the latter when the underlying distributions are heavy-tailed. We apply our method to EEG recordings from a virtual-reality driving experiment and to calcium imaging recordings in inhibitory and excitatory neurons of the auditory cortex in mice subjected to sound stimuli. Our findings reveal distinct regional dependencies across frequency bands and brain states.

stat.ME

Classification of High-dimensional Time Series in Spectral Domain using Explainable Features

Interpretable classification of time series presents significant challenges in high dimensions. Traditional feature selection methods in the frequency domain often assume sparsity in spectral density matrices (SDMs) or their inverses, which can be restrictive for real-world applications. In this article, we propose a model-based approach for classifying high-dimensional stationary time series by assuming sparsity in the difference between inverse SDMs. Our approach emphasizes the interpretability of model parameters, making it especially suitable for fields like neuroscience, where understanding differences in brain network connectivity across various states is crucial. The estimators for model parameters demonstrate consistency under appropriate conditions. We further propose using standard deep learning optimizers for parameter estimation, employing techniques such as mini-batching and learning rate scheduling. Additionally, we introduce a method to screen the most discriminatory frequencies for classification, which exhibits the sure screening property under general conditions. The flexibility of the proposed model allows the significance of covariates to vary across frequencies, enabling nuanced inferences and deeper insights into the underlying problem. The novelty of our method lies in the interpretability of the model parameters, addressing critical needs in neuroscience. The proposed approaches have been evaluated on simulated examples and the `Alert-vs-Drowsy' EEG dataset.

stat.ML

Robust Classification of High-Dimensional Data using Data-Adaptive Energy Distance

Classification of high-dimensional low sample size (HDLSS) data poses a challenge in a variety of real-world situations, such as gene expression studies, cancer research, and medical imaging. This article presents the development and analysis of some classifiers that are specifically designed for HDLSS data. These classifiers are free of tuning parameters and are robust, in the sense that they are devoid of any moment conditions of the underlying data distributions. It is shown that they yield perfect classification in the HDLSS asymptotic regime, under some fairly general conditions. The comparative performance of the proposed classifiers is also investigated. Our theoretical results are supported by extensive simulation studies and real data analysis, which demonstrate promising advantages of the proposed classification techniques over several widely recognized methods.

stat.ML

On Exact Feature Screening in Ultrahigh-dimensional Binary Classification

We propose a new model-free feature screening method based on energy distances for ultrahigh-dimensional binary classification problems. With a high probability, the proposed method retains only relevant features after discarding all the noise variables. The proposed screening method is also extended to identify pairs of variables that are marginally undetectable but have differences in their joint distributions. Finally, we build a classifier that maintains coherence between the proposed feature selection criteria and discrimination method and also establish its risk consistency. An extensive numerical study with simulated and real benchmark data sets shows clear and convincing advantages of our proposed method over the state-of-the-art methods.

stat.ME

On Generalizations of Some Distance Based Classifiers for HDLSS Data

In high dimension, low sample size (HDLSS) settings, classifiers based on Euclidean distances like the nearest neighbor classifier and the average distance classifier perform quite poorly if differences between locations of the underlying populations get masked by scale differences. To rectify this problem, several modifications of these classifiers have been proposed in the literature. However, existing methods are confined to location and scale differences only, and often fail to discriminate among populations differing outside of the first two moments. In this article, we propose some simple transformations of these classifiers resulting into improved performance even when the underlying populations have the same location and scale. We further propose a generalization of these classifiers based on the idea of grouping of variables. The high-dimensional behavior of the proposed classifiers is studied theoretically. Numerical experiments with a variety of simulated examples as well as an extensive analysis of real data sets exhibit advantages of the proposed methods.

stat.ME

On a Generalization of the Average Distance Classifier

In high dimension, low sample size (HDLSS)settings, the simple average distance classifier based on the Euclidean distance performs poorly if differences between the locations get masked by the scale differences. To rectify this issue, modifications to the average distance classifier was proposed by Chan and Hall (2009). However, the existing classifiers cannot discriminate when the populations differ in other aspects than locations and scales. In this article, we propose some simple transformations of the average distance classifier to tackle this issue. The resulting classifiers perform quite well even when the underlying populations have the same location and scale. The high-dimensional behaviour of the proposed classifiers is studied theoretically. Numerical experiments with a variety of simulated as well as real data sets exhibit the usefulness of the proposed methodology.

stat.ME