Searcharxiv⌕ Search

arXiv subjects

Muhammad Sudipto Siam Dip

Publications and source records attributed to Muhammad Sudipto Siam Dip.

4 recordsLinked to original sources

SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations

Test-time adaptation (TTA) addresses distribution shift using only unlabeled test data. Existing methods typically adapt pretrained models by updating their parameters, limiting both what is adapted and where adaptation can occur within the network. We instead keep the pretrained network frozen and steer its intermediate representations. We introduce SPeaR (Steering Primitive for Realigning Representations), which inserts lightweight learnable modules at stage boundaries and optimizes them directly from the test stream, requiring neither source data nor supervised warm-up. Each primitive is optimized using a gated objective that reduces uncertainty only when adaptation is beneficial, along with a diversity regularizer to prevent collapse, and a multi-depth anchor to stabilize adaptation. We show that steering early representations is the most effective strategy, and that the same primitive transfers across convolutional and Transformer architectures. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-C, SPeaR consistently matches or outperforms methods that adapt orders of magnitude more parameters, remains robust across a wide range of batch sizes, and preserves source-domain performance during continual adaptation.

cs.LG↗

NeuroSleepNet: An Explainable Multi-Head Attention-Based Framework With Spatial and Multi-Scale Independent Temporal Context Learning for Automatic Sleep Stage Scoring

Automatic sleep stage scoring is crucial for monitoring sleep quality and sleep-related disorders. However, most existing electroencephalogram (EEG) based deep learning frameworks rely on current and previous input sequences (varying from 4-100 epochs), which increases computational complexity and creates dependency on transition rules. In this study, we propose NeuroSleepNet, a deep-learning framework that classifies sleep stages using only the current input sequence. NeuroSleepNet employs a two-stage representation learning that includes spatial and multi-scale independent temporal context learning. Multi-head self-attention transformer encoders were employed to transform extracted representations into contextual embeddings to improve the distinctiveness of learned features. Additionally, to handle the class imbalances in different sleep stages, a logarithmic scale-based weighting technique was incorporated into the loss function of NeuroSleepNet. Furthermore, an explainability module was integrated into NeuroSleepNet to visualize and interpret class-specific decision patterns. NeuroSleepNet demonstrated state-of-the-art performance, achieving an accuracy of 86.1 percent, a macro-F1 score of 80.8 percent, and a Cohen's kappa of 0.805 on the Sleep-EDF expanded dataset. It achieved 82.0 percent, 76.3 percent, and 0.753 on MESA, 80.5 percent, 76.8 percent, and 0.738 on Physio2018, and 86.7 percent, 80.9 percent, and 0.804 on the SHHS database, respectively, by utilizing only the current sequence. As a result, there is no need for previous long sequences or sequential dependencies, making it computationally efficient and well-suited for long-term automatic sleep monitoring.

eess.SP↗

Optimized Feature Selection and Neural Network-Based Classification of Motor Imagery Using EEG Signals

Objective: Machine learning- and deep learning-based models have recently been employed in motor imagery intention classification from electroencephalogram (EEG) signals. Nevertheless, there is a limited understanding of feature selection to assist in identifying the most significant features in different spatial locations. Methods: This study proposes a feature selection technique using sequential forward feature selection with support vector machines and feeding the selected features to deep neural networks to classify motor imagery intention using multi-channel EEG. Results: The proposed model was evaluated with a publicly available dataset and achieved an average accuracy of 79.70 percent with a standard deviation of 7.98 percent for classifying two motor imagery scenarios. Conclusions: These results demonstrate that our method effectively identifies the most informative and discriminative characteristics of neural activity at different spatial locations, offering potential for future prosthetics and brain-computer interface applications. Significance: This approach enhances model performance while identifying key spatial EEG features, advancing brain-computer interfaces and prosthetic systems.

eess.SP↗

oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models

In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper

eess.AS↗