SearcharxivSearch

arXiv subjects

Neeraj Kumar

Publications and source records attributed to Neeraj Kumar.

At least 37 records · Page 2Linked to original sources

Electrical Load Forecasting in Smart Grid: A Personalized Federated Learning Approach

Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) are used to record household energy consumption. Traditional machine learning (ML) methods are often employed for load forecasting but require data sharing which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. This paper presents a novel personalized federated learning (PFL) method to load prediction under non-independent and identically distributed (non-IID) metering data settings. Specifically, we introduce meta-learning, where the learning rates are manipulated using the meta-learning idea to maximize the gradient for each client in each global round. Clients with varying processing capacities, data sizes, and batch sizes can participate in global model aggregation and improve their local load forecasting via personalized learning. Simulation results show that our approach outperforms state-of-the-art ML and FL methods in terms of better load forecasting accuracy.

cs.LG

Emergence of interfacial magnetism in strongly-correlated nickelate-titanate superlattices

Strongly-correlated transition-metal oxides are widely known for their various exotic phenomena. This is exemplified by rare-earth nickelates such as LaNiO$_{3}$, which possess intimate interconnections between their electronic, spin, and lattice degrees of freedom. Their properties can be further enhanced by pairing them in hybrid heterostructures, which can lead to hidden phases and emergent phenomena. An important example is the LaNiO$_{3}$/LaTiO$_{3}$ superlattice, where an interlayer electron transfer has been observed from LaTiO$_{3}$ into LaNiO$_{3}$ leading to a high-spin state. However, macroscopic emergence of magnetic order associated with this high-spin state has so far not been observed. Here, by using muon spin rotation, x-ray absorption, and resonant inelastic x-ray scattering, we present direct evidence of an emergent antiferromagnetic order with high magnon energy and exchange interactions at the LaNiO$_{3}$/LaTiO$_{3}$ interface. As the magnetism is purely interfacial, a single LaNiO$_{3}$/LaTiO$_{3}$ interface can essentially behave as an atomically thin strongly-correlated quasi-two-dimensional antiferromagnet, potentially allowing its technological utilisation in advanced spintronic devices. Furthermore, its strong quasi-two-dimensional magnetic correlations, orbitally-polarized planar ligand holes, and layered superlattice design make its electronic, magnetic, and lattice configurations resemble the precursor states of superconducting cuprates and nickelates, but with an $S \rightarrow 1$ spin state instead.

cond-mat.str-el

Learning-based Two-tiered Online Optimization of Region-wide Datacenter Resource Allocation

Online optimization of resource management for large-scale data centers and infrastructures to meet dynamic capacity reservation demands and various practical constraints (e.g., feasibility and robustness) is a very challenging problem. Mixed Integer Programming (MIP) approaches suffer from recognized limitations in such a dynamic environment, while learning-based approaches may face with prohibitively large state/action spaces. To this end, this paper presents a novel two-tiered online optimization to enable a learning-based Resource Allowance System (RAS). To solve optimal server-to-reservation assignment in RAS in an online fashion, the proposed solution leverages a reinforcement learning (RL) agent to make high-level decisions, e.g., how much resource to select from the Main Switch Boards (MSBs), and then a low-level Mixed Integer Linear Programming (MILP) solver to generate the local server-to-reservation mapping, conditioned on the RL decisions. We take into account fault tolerance, server movement minimization, and network affinity requirements and apply the proposed solution to large-scale RAS problems. To provide interpretability, we further train a decision tree model to explain the learned policies and to prune unreasonable corner cases at the low-level MILP solver, resulting in further performance improvement. Extensive evaluations show that our two-tiered solution outperforms baselines such as pure MIP solver by over $15\%$ while delivering $100\times$ speedup in computation.

cs.NI

Assessing Reusability of Deep Learning-Based Monotherapy Drug Response Prediction Models Trained with Omics Data

Cancer drug response prediction (DRP) models present a promising approach towards precision oncology, tailoring treatments to individual patient profiles. While deep learning (DL) methods have shown great potential in this area, models that can be successfully translated into clinical practice and shed light on the molecular mechanisms underlying treatment response will likely emerge from collaborative research efforts. This highlights the need for reusable and adaptable models that can be improved and tested by the wider scientific community. In this study, we present a scoring system for assessing the reusability of prediction DRP models, and apply it to 17 peer-reviewed DL-based DRP models. As part of the IMPROVE (Innovative Methodologies and New Data for Predictive Oncology Model Evaluation) project, which aims to develop methods for systematic evaluation and comparison DL models across scientific domains, we analyzed these 17 DRP models focusing on three key categories: software environment, code modularity, and data availability and preprocessing. While not the primary focus, we also attempted to reproduce key performance metrics to verify model behavior and adaptability. Our assessment of 17 DRP models reveals both strengths and shortcomings in model reusability. To promote rigorous practices and open-source sharing, we offer recommendations for developing and sharing prediction models. Following these recommendations can address many of the issues identified in this study, improving model reusability without adding significant burdens on researchers. This work offers the first comprehensive assessment of reusability and reproducibility across diverse DRP models, providing insights into current model sharing practices and promoting standards within the DRP and broader AI-enabled scientific research community.

q-bio.BM

CACTUS: Chemistry Agent Connecting Tool-Usage to Science

Large language models (LLMs) have shown remarkable potential in various domains, but they often lack the ability to access and reason over domain-specific knowledge and tools. In this paper, we introduced CACTUS (Chemistry Agent Connecting Tool-Usage to Science), an LLM-based agent that integrates cheminformatics tools to enable advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama2-7b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b and Mistral-7b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with domain-specific tools, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment. Furthermore, CACTUS represents a significant milestone in the field of cheminformatics, offering an adaptable tool for researchers engaged in chemistry and molecular discovery. By integrating the strengths of open-source LLMs with domain-specific tools, CACTUS has the potential to accelerate scientific advancement and unlock new frontiers in the exploration of novel, effective, and safe therapeutic candidates, catalysts, and materials. Moreover, CACTUS's ability to integrate with automated experimentation platforms and make data-driven decisions in real time opens up new possibilities for autonomous discovery.

cs.CL

Regularity of powers of d-sequence (parity) binomial edge ideals of unicycle graphs

We classify all unicycle graphs whose edge-binomials form a $d$-sequence, particularly linear type binomial edge ideals. We also classify unicycle graphs whose parity edge-binomials form a $d$-sequence. We study the regularity of powers of (parity) binomial edge ideals of unicycle graphs generated by $d$-sequence (parity) edge-binomials.

math.AC

Scaffold-Based Multi-Objective Drug Candidate Optimization

In therapeutic design, balancing various physiochemical properties is crucial for molecule development, similar to how Multiparameter Optimization (MPO) evaluates multiple variables to meet a primary goal. While many molecular features can now be predicted using \textit{in silico} methods, aiding early drug development, the vast data generated from high throughput virtual screening challenges the practicality of traditional MPO approaches. Addressing this, we introduce a scaffold focused graph-based Markov chain Monte Carlo framework (ScaMARS) built to generate molecules with optimal properties. This innovative framework is capable of self-training and handling a wider array of properties, sampling different chemical spaces according to the starting scaffold. The benchmark analysis on several properties shows that ScaMARS has a diversity score of 84.6\% and has a much higher success rate of 99.5\% compared to conditional models. The integration of new features into MPO significantly enhances its adaptability and effectiveness in therapeutic design, facilitating the discovery of candidates that efficiently optimize multiple properties.

q-bio.BM

Style Description based Text-to-Speech with Conditional Prosodic Layer Normalization based Diffusion GAN

In this paper, we present a Diffusion GAN based approach (Prosodic Diff-TTS) to generate the corresponding high-fidelity speech based on the style description and content text as an input to generate speech samples within only 4 denoising steps. It leverages the novel conditional prosodic layer normalization to incorporate the style embeddings into the multi head attention based phoneme encoder and mel spectrogram decoder based generator architecture to generate the speech. The style embedding is generated by fine tuning the pretrained BERT model on auxiliary tasks such as pitch, speaking speed, emotion,gender classifications. We demonstrate the efficacy of our proposed architecture on multi-speaker LibriTTS and PromptSpeech datasets, using multiple quantitative metrics that measure generated accuracy and MOS.

cs.SD

Rees algebra of maximal order Pfaffians and its diagonal subalgebras

Given a skew-symmetric matrix $X$, the Pfaffian of $X$ is defined as the square root of the determinant of $X$. In this article, we give the explicit defining equations of the Rees algebra of a Pfaffian ideal $I$ generated by the maximal order Pfaffians of a generic skew-symmetric matrix. We further prove that all diagonal subalgebras of the corresponding Rees algebra of $I$ are Koszul. We also look at Rees algebras of Pfaffian ideals of linear type associated with certain sparse skew-symmetric matrices. In particular, we consider the tridiagonal matrices and identify the corresponding Pfaffian ideals to be of Gröbner linear type and as the vertex cover ideals of unmixed bipartite graphs. As an application of our results, we conclude that all their ordinary and symbolic powers have linear quotients.

math.AC

Distraction-free Embeddings for Robust VQA

The generation of effective latent representations and their subsequent refinement to incorporate precise information is an essential prerequisite for Vision-Language Understanding (VLU) tasks such as Video Question Answering (VQA). However, most existing methods for VLU focus on sparsely sampling or fine-graining the input information (e.g., sampling a sparse set of frames or text tokens), or adding external knowledge. We present a novel "DRAX: Distraction Removal and Attended Cross-Alignment" method to rid our cross-modal representations of distractors in the latent space. We do not exclusively confine the perception of any input information from various modalities but instead use an attention-guided distraction removal method to increase focus on task-relevant information in latent embeddings. DRAX also ensures semantic alignment of embeddings during cross-modal fusions. We evaluate our approach on a challenging benchmark (SUTD-TrafficQA dataset), testing the framework's abilities for feature and event queries, temporal relation understanding, forecasting, hypothesis, and causal analysis through extensive experiments.

cs.CV

Charge fluctuations in the intermediate-valence ground state of SmCoIn$_5$

The microscopic mechanism of heavy band formation, relevant for unconventional superconductivity in CeCoIn$_5$ and other Ce-based heavy fermion materials, depends strongly on the efficiency with which $f$ electrons are delocalized from the rare earth sites and participate in a Kondo lattice. Replacing Ce$^{3+}$ ($4f^1$, $J=5/2$) with Sm$^{3+}$ ($4f^5$, $J=5/2$), we show that a combination of crystal field and on-site Coulomb repulsion causes SmCoIn$_5$ to exhibit a $Γ_7$ ground state similar to CeCoIn$_5$ with multiple $f$ electrons. Remarkably, we also find that with this ground state, SmCoIn$_5$ exhibits a temperature-induced valence crossover consistent with a Kondo scenario, leading to increased delocalization of $f$ holes below a temperature scale set by the crystal field, $T_v$ $\approx$ 60 K. Our result provides evidence that in the case of many $f$ electrons, the crystal field remains the most important tuning knob in controlling the efficiency of delocalization near a heavy fermion quantum critical point, and additionally clarifies that charge fluctuations play a general role in the ground state of "115" materials.

cond-mat.str-el

An Effective Meaningful Way to Evaluate Survival Models

One straightforward metric to evaluate a survival prediction model is based on the Mean Absolute Error (MAE) -- the average of the absolute difference between the time predicted by the model and the true event time, over all subjects. Unfortunately, this is challenging because, in practice, the test set includes (right) censored individuals, meaning we do not know when a censored individual actually experienced the event. In this paper, we explore various metrics to estimate MAE for survival datasets that include (many) censored individuals. Moreover, we introduce a novel and effective approach for generating realistic semi-synthetic survival datasets to facilitate the evaluation of metrics. Our findings, based on the analysis of the semi-synthetic datasets, reveal that our proposed metric (MAE using pseudo-observations) is able to rank models accurately based on their performance, and often closely matches the true MAE -- in particular, is better than several alternative methods.

cs.LG

The ACROBAT 2022 Challenge: Automatic Registration Of Breast Cancer Tissue

The alignment of tissue between histopathological whole-slide-images (WSI) is crucial for research and clinical applications. Advances in computing, deep learning, and availability of large WSI datasets have revolutionised WSI analysis. Therefore, the current state-of-the-art in WSI registration is unclear. To address this, we conducted the ACROBAT challenge, based on the largest WSI registration dataset to date, including 4,212 WSIs from 1,152 breast cancer patients. The challenge objective was to align WSIs of tissue that was stained with routine diagnostic immunohistochemistry to its H&E-stained counterpart. We compare the performance of eight WSI registration algorithms, including an investigation of the impact of different WSI properties and clinical covariates. We find that conceptually distinct WSI registration methods can lead to highly accurate registration performances and identify covariates that impact performances across methods. These results establish the current state-of-the-art in WSI registration and guide researchers in selecting and developing methods.

eess.IV

Breaking of universal nature of central charge criticality in $AdS$ black holes in Gauss-Bonnet Gravity

In this paper, we have studied the thermodynamics of Gauss-Bonnet black holes in D-dimensional $AdS$ spacetime. Here, the cosmological constant ($Λ$), Newton's gravitational constant ($G$) and the Gauss-Bonnet parameter ($α$) are varied in the bulk, and a mixed first law is rewritten considering central charge ($C$) (of dual boundary conformal theory) and its conjugate variable utilising the gauge-gravity duality. A novel universal nature of central charge near the critical point of black hole phase transition in Einstein's gravity has been observed in \cite{mann1}. We observe that this universal nature breaks when such phase transition is considered for black holes in the Gauss-Bonnet gravity. Apart from this, treating the Gauss-Bonnet parameter as a thermodynamic variable as suggested in \cite{kastori} in light of the consistency between first law and the Smarr relation leads to modified thermodynamic volume (conjugate to variable cosmological constant), adding to a new understanding of the Van der Waals gas like behaviour of the black holes in higher dimensional and higher curvature gravity theories. Our analysis considers a general $D$ dimensional background. We have then imposed a greater focus in the analysis of the phase structure of the five dimensional Gauss-Bonnet spacetime. Our analysis also shows that the general universal nature of the critical value of the central charge (which was present in four dimensional $AdS$ spacetime), breaks down in case of five dimensional $AdS$ spacetime even in the absence of Gauss-Bonnet gravity. This finding indicates the universal nature of the central charge may be a special feature of the four dimensional $AdS$ spacetime only.

gr-qc

KL Regularized Normalization Framework for Low Resource Tasks

Large pre-trained models, such as Bert, GPT, and Wav2Vec, have demonstrated great potential for learning representations that are transferable to a wide variety of downstream tasks . It is difficult to obtain a large quantity of supervised data due to the limited availability of resources and time. In light of this, a significant amount of research has been conducted in the area of adopting large pre-trained datasets for diverse downstream tasks via fine tuning, linear probing, or prompt tuning in low resource settings. Normalization techniques are essential for accelerating training and improving the generalization of deep neural networks and have been successfully used in a wide variety of applications. A lot of normalization techniques have been proposed but the success of normalization in low resource downstream NLP and speech tasks is limited. One of the reasons is the inability to capture expressiveness by rescaling parameters of normalization. We propose KullbackLeibler(KL) Regularized normalization (KL-Norm) which make the normalized data well behaved and helps in better generalization as it reduces over-fitting, generalises well on out of domain distributions and removes irrelevant biases and features with negligible increase in model parameters and memory overheads. Detailed experimental evaluation on multiple low resource NLP and speech tasks, demonstrates the superior performance of KL-Norm as compared to other popular normalization and regularization techniques.

cs.CL

Dynamic Molecular Graph-based Implementation for Biophysical Properties Prediction

Neural Networks (GNNs) have revolutionized the molecular discovery to understand patterns and identify unknown features that can aid in predicting biophysical properties and protein-ligand interactions. However, current models typically rely on 2-dimensional molecular representations as input, and while utilization of 2\3- dimensional structural data has gained deserved traction in recent years as many of these models are still limited to static graph representations. We propose a novel approach based on the transformer model utilizing GNNs for characterizing dynamic features of protein-ligand interactions. Our message passing transformer pre-trains on a set of molecular dynamic data based off of physics-based simulations to learn coordinate construction and make binding probability and affinity predictions as a downstream task. Through extensive testing we compare our results with the existing models, our MDA-PLI model was able to outperform the molecular interaction prediction models with an RMSE of 1.2958. The geometric encodings enabled by our transformer architecture and the addition of time series data add a new dimensionality to this form of research.

cs.LG

Swarm of UAVs for Network Management in 6G: A Technical Review

Fifth-generation (5G) cellular networks have led to the implementation of beyond 5G (B5G) networks, which are capable of incorporating autonomous services to swarm of unmanned aerial vehicles (UAVs). They provide capacity expansion strategies to address massive connectivity issues and guarantee ultra-high throughput and low latency, especially in extreme or emergency situations where network density, bandwidth, and traffic patterns fluctuate. On the one hand, 6G technology integrates AI/ML, IoT, and blockchain to establish ultra-reliable, intelligent, secure, and ubiquitous UAV networks. 6G networks, on the other hand, rely on new enabling technologies such as air interface and transmission technologies, as well as a unique network design, posing new challenges for the swarm of UAVs. Keeping these challenges in mind, this article focuses on the security and privacy, intelligence, and energy-efficiency issues faced by swarms of UAVs operating in 6G mobile networks. In this state-of-the-art review, we integrated blockchain and AI/ML with UAV networks utilizing the 6G ecosystem. The key findings are then presented, and potential research challenges are identified. We conclude the review by shedding light on future research in this emerging field of research.

cs.NI