SearcharxivSearch

arXiv subjects

Sahil Kumar

Publications and source records attributed to Sahil Kumar.

7 recordsLinked to original sources

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

MambaVoiceCloning (MVC) asks whether the conditioning path of diffusion-based TTS can be made fully SSM-only at inference, removing all attention and explicit RNN-style recurrence layers across text, rhythm, and prosody, while preserving or improving quality under controlled conditions. MVC combines a gated bidirectional Mamba text encoder, a Temporal Bi-Mamba supervised by a lightweight alignment teacher discarded after training, and an Expressive Mamba with AdaLN modulation, yielding linear-time O(T) conditioning with bounded activation memory and practical finite look-ahead streaming. Unlike prior Mamba-TTS systems that remain hybrid at inference, MVC removes attention-based duration and style modules under a fixed StyleTTS2 mel-diffusion-vocoder backbone. Trained on LJSpeech/LibriTTS and evaluated on VCTK, CSS10 (ES/DE/FR), and long-form Gutenberg passages, MVC achieves modest but statistically reliable gains over StyleTTS2, VITS, and Mamba-attention hybrids in MOS/CMOS, F0 RMSE, MCD, and WER, while reducing encoder parameters to 21M and improving throughput by 1.6x. Diffusion remains the dominant latency source, but SSM-only conditioning improves memory footprint, stability, and deployability.

cs.SD

A Decomposable Forward Process in Diffusion Models for Time-Series Forecasting

We introduce a model-agnostic forward diffusion process for time-series forecasting that decomposes signals into spectral components, preserving structured temporal patterns such as seasonality more effectively than standard diffusion. Unlike prior work that modifies the network architecture or diffuses directly in the frequency domain, our proposed method alters only the diffusion process itself, making it compatible with existing diffusion backbones (e.g., DiffWave, TimeGrad, CSDI). By staging noise injection according to component energy, it maintains high signal-to-noise ratios for dominant frequencies throughout the diffusion trajectory, thereby improving the recoverability of long-term patterns. This strategy enables the model to maintain the signal structure for a longer period in the forward process, leading to improved forecast quality. Across standard forecasting benchmarks, we show that applying spectral decomposition strategies, such as the Fourier or Wavelet transform, consistently improves upon diffusion models using the baseline forward process, with negligible computational overhead. The code for this paper is available at https://anonymous.4open.science/r/D-FDP-4A29.

stat.ML

Unravelling the Catalytic Activity of Dual-Metal Doped N6-Graphene for Sulfur Reduction via Machine Learning-Accelerated First-Principles Calculations

Understanding and optimizing polysulfide adsorption and conversion processes are critical to mitigating shuttle effects and sluggish redox kinetics in lithium-sulfur batteries (LSBs). Here, we introduce a machine-learning-accelerated framework, Precise and Accurate Configuration Evaluation (PACE), that integrates Machine Learning Interatomic Potentials (MLIPs) with Density Functional Theory (DFT) to systematically explore adsorption configurations and energetics of a series of N6-coordinated dual-atom catalysts (DACs). Our results demonstrate that, compared with single-atom catalysts, DACs exhibit improved LiPS adsorption and redox conversion through cooperative metal-sulfur interactions and electronic coupling between adjacent metal centers. Among all DACs, Fe-Ni and Fe-Pt show optimal catalytic performance, due to their optimal adsorption energies (-1.0 to -2.3 eV), low free-energy barriers (<=0.4 eV) for the Li2S2 to Li2S conversion, and facile Li2S decomposition barriers (<=1.0 eV). To accelerate catalyst screening, we further developed a machine learning (ML) regression model trained on DFT-calculated data to predict the Gibbs free energy (\Delta G) of Li2Sn adsorption using physically interpretable descriptors. The Gradient Boosting Regression (GBR) model yields an R^2 of 0.85 and an MAE of 0.26 eV, enabling the rapid prediction of \Delta G for unexplored DACs. Electronic-structure analyses reveal that the superior performance originates from the optimal d-band alignment and S-S bond polarization induced by the cooperative effect of dual metal centres. This dual ML-DFT framework demonstrates a generalizable, data-driven design strategy for the rational discovery of efficient catalysts for next-generation LSBs.

cond-mat.mtrl-sci

KatzBot: Revolutionizing Academic Chatbot for Enhanced Communication

Effective communication within universities is crucial for addressing the diverse information needs of students, alumni, and external stakeholders. However, existing chatbot systems often fail to deliver accurate, context-specific responses, resulting in poor user experiences. In this paper, we present KatzBot, an innovative chatbot powered by KatzGPT, a custom Large Language Model (LLM) fine-tuned on domain-specific academic data. KatzGPT is trained on two university-specific datasets: 6,280 sentence-completion pairs and 7,330 question-answer pairs. KatzBot outperforms established existing open source LLMs, achieving higher accuracy and domain relevance. KatzBot offers a user-friendly interface, significantly enhancing user satisfaction in real-world applications. The source code is publicly available at \url{https://github.com/AiAI-99/katzbot}.

cs.CL

Vision Transformer Segmentation for Visual Bird Sound Denoising

Audio denoising, especially in the context of bird sounds, remains a challenging task due to persistent residual noise. Traditional and deep learning methods often struggle with artificial or low-frequency noise. In this work, we propose ViTVS, a novel approach that leverages the power of the vision transformer (ViT) architecture. ViTVS adeptly combines segmentation techniques to disentangle clean audio from complex signal mixtures. Our key contributions encompass the development of ViTVS, introducing comprehensive, long-range, and multi-scale representations. These contributions directly tackle the limitations inherent in conventional approaches. Extensive experiments demonstrate that ViTVS outperforms state-of-the-art methods, positioning it as a benchmark solution for real-world bird sound denoising applications. Source code is available at: https://github.com/aiai-4/ViVTS.

cs.SD

Comparative Study of MPPT and Parameter Estimation of PV cells

The presented work focuses on utilising machine learning techniques to accurately estimate accurate values for known and unknown parameters of the PVLIB model for solar cells and photovoltaic modules.Finding accurate model parameters of circuits for photovoltaic (PV) cells is important for a variety of tasks. An Artificial Neural Network (ANN) algorithm was employed, which outperformed other metaheuristic and machine learning algorithms in terms of computational efficiency. To validate the consistency of the data and output, the results were compared against other machine learning algorithms based on irradiance and temperature. A Bland Altman test was conducted that resulted in more than 95 percent accuracy rate. Upon validation, the ANN algorithm was utilised to estimate the parameters and their respective values.

cs.LG

Leptogenesis and Neutrinoless Double Beta Decay in the Scotogenic Hybrid Textures of Neutrino Mass Matrix

In our recent work we identify the hybrid textures of neutrino mass matrix which simultaneously account for dark matter (DM) and neutrinoless double beta decay ($0\nu\beta\beta$). We also obtained the bounds on dark matter mass and effective Majorana mass $|M_{ee}|$. In this work we look for those hybrid textures which altogether accounts for DM, $0\nu\beta\beta$ and leptogenesis. We have found correlation of baryon asymmetry of universe $Y$ with dark matter mass $M_1$ and effective Majorana mass $|M_{ee}|$. We use experimental bounds on relic density of dark matter ($\Omega h^2$) and baryon asymmetry of universe to identify the hybrid textures. We found that out of five hybrid textures which simultaneously satisfies the physics observations of the DM and $0\nu\beta\beta$ only three hybrid textures altogether satisfy the DM, $0\nu\beta\beta$ and leptogenesis. It is interesting to note that these three hybrid textures gives lower bound to the effective Majorana mass $|M_{ee}|$ which can be probed in current and future experiments like SuperNEMO, KamLAND-Zen, NEXT, and nEXO (5 year) have sensitivity reaches of 0.05 eV, 0.045 eV, 0.03 eV, and 0.015 eV, respectively.

hep-ph