SearcharxivSearch

arXiv subjects

Neeraj Sharma

Publications and source records attributed to Neeraj Sharma.

10 recordsLinked to original sources

Knowledge-guided Transfer Prediction In Underrepresented Populations: A GRU-D-Static Framework For Maternal And Neonatal Outcomes

Integrating summary-level scientific knowledge into neural network models provides a practical strategy for transferring prediction models trained on adequately sampled source cohorts to underrepresented target populations, where individual-level data in the target domain are often limited or unavailable. In this study, we propose transfer prediction strategies incorporating external summary-level scientific knowledge and illustrate its application on the PRISMA Maternal and Neonatal Health Study, training a neural network model on the source data to predict adverse outcomes in the target cohorts. Besides, we also extend the existing GRU-D framework by incorporating static feature embeddings and attention weights to jointly leverage temporal and static information for improved prediction. Our approach employs soft labels derived from summary-level statistics describing the target population to fine-tune GRU-D-Static models that are initially trained on the source populations. We evaluate six maternal and neonatal outcomes, including stillbirth, preterm birth, low birth weight, small vulnerable newborn, neonatal death, and maternal near miss. Across all tested scenarios, fine-tuning using soft labels from just basic covariates substantially improved predictive performance compared with deep learning models trained on the source sample. Furthermore, the performance slightly improves more when additional covariates were incorporated into the logistic regression model or when partial input features from the target population were available for fine-tuning. These findings demonstrate that integrating existing scientific knowledge in the literature through transfer prediction of source neural network models can enhance prediction performance in underrepresented target populations, reducing reliance on large-scale data collection and supporting risk prediction in global health.

stat.AP

Optimal Thermalization under Indefinite Causal Order with Identical and Asymmetric Baths

Indefinite causal order (ICO), in which the order of quantum operations is placed in a coherent superposition, has been demonstrated to enhance various information-processing tasks. Here, we investigate its impact on the thermodynamic processes generated by thermalizing quantum channels. We consider a two-level system interacting with two thermal baths under a quantum SWITCH, with the channel order controlled coherently by an ancillary qubit. We derive closed-form expressions for the effective inverse temperature $\beta_f$ of the postselected system state for both identical and distinct bath temperatures, and identify the control-qubit parameters that maximize heating or cooling. Our analysis reveals how the diagonal and coherent components of the control-qubit state contribute separately to the temperature shift, and how their interplay enables departures from the thermal response attainable under protocols with a definite causal order within the thermodynamic setting considered here. Bath asymmetry enhances these effects, while reduced purity of the control qubit state suppresses them. These results provide a systematic framework for assessing SWITCH-based thermalization in the setting of indefinite causal order, and identify control-qubit coherence as a tunable resource.

quant-ph

From Time and Place to Preference: LLM-Driven Geo-Temporal Context in Recommendations

Most recommender systems treat timestamps as numeric or cyclical values, overlooking real-world context such as holidays, events, and seasonal patterns. We propose a scalable framework that uses large language models (LLMs) to generate geo-temporal embeddings from only a timestamp and coarse location, capturing holidays, seasonal trends, and local/global events. We then introduce a geo-temporal embedding informativeness test as a lightweight diagnostic, demonstrating on MovieLens, LastFM, and a production dataset that these embeddings provide predictive signal consistent with the outcomes of full model integrations. Geo-temporal embeddings are incorporated into sequential models through (1) direct feature fusion with metadata embeddings or (2) an auxiliary loss that enforces semantic and geo-temporal alignment. Our findings highlight the need for adaptive or hybrid recommendation strategies, and we release a context-enriched MovieLens dataset to support future research.

cs.IR

Predicting Movie Hits Before They Happen with LLMs

Addressing the cold-start issue in content recommendation remains a critical ongoing challenge. In this work, we focus on tackling the cold-start problem for movies on a large entertainment platform. Our primary goal is to forecast the popularity of cold-start movies using Large Language Models (LLMs) leveraging movie metadata. This method could be integrated into retrieval systems within the personalization pipeline or could be adopted as a tool for editorial teams to ensure fair promotion of potentially overlooked movies that may be missed by traditional or algorithmic solutions. Our study validates the effectiveness of this approach compared to established baselines and those we developed.

cs.IR

Monitoring lead-acid battery function using operando neutron radiography

Investigating batteries while they operate allows researchers to track the inner electrochemical processes involved in working conditions. This study describes the first neutron radiography investigation of a lead-acid battery. A custom-designed neutron friendly lead-acid cell and casing is developed and studied operando during electrochemical cycling, in order to observe the activity within the electrolyte and at the electrodes. This experimental work is coupled with Monte Carlo simulations of neutron transmittance. Details of cell construction, data collection and data analysis are presented. This work highlights the potential of neutron imaging for tracking battery function and outlines opportunities for further development.

physics.ins-det

Multi-modal Point-of-Care Diagnostics for COVID-19 Based On Acoustics and Symptoms

The research direction of identifying acoustic bio-markers of respiratory diseases has received renewed interest following the onset of COVID-19 pandemic. In this paper, we design an approach to COVID-19 diagnostic using crowd-sourced multi-modal data. The data resource, consisting of acoustic signals like cough, breathing, and speech signals, along with the data of symptoms, are recorded using a web-application over a period of ten months. We investigate the use of statistical descriptors of simple time-frequency features for acoustic signals and binary features for the presence of symptoms. Unlike previous works, we primarily focus on the application of simple linear classifiers like logistic regression and support vector machines for acoustic data while decision tree models are employed on the symptoms data. We show that a multi-modal integration of acoustics and symptoms classifiers achieves an area-under-curve (AUC) of 92.40, a significant improvement over any individual modality. Several ablation experiments are also provided which highlight the acoustic and symptom dimensions that are important for the task of COVID-19 diagnostics.

eess.AS

DiCOVA Challenge: Dataset, task, and baseline system for COVID-19 diagnosis using acoustics

The DiCOVA challenge aims at accelerating research in diagnosing COVID-19 using acoustics (DiCOVA), a topic at the intersection of speech and audio processing, respiratory health diagnosis, and machine learning. This challenge is an open call for researchers to analyze a dataset of sound recordings collected from COVID-19 infected and non-COVID-19 individuals for a two-class classification. These recordings were collected via crowdsourcing from multiple countries, through a website application. The challenge features two tracks, one focusing on cough sounds, and the other on using a collection of breath, sustained vowel phonation, and number counting speech recordings. In this paper, we introduce the challenge and provide a detailed description of the task, and present a baseline system for the task.

eess.AS

Role of Attentive History Selection in Conversational Information Seeking

The rise of intelligent assistant systems like Siri and Alexa have led to the emergence of Conversational Search, a research track of Information Retrieval (IR) that involves interactive and iterative information-seeking user-system dialog. Recently released OR-QuAC and TCAsT19 datasets narrow their research focus on the retrieval aspect of conversational search i.e. fetching the relevant documents (passages) from a large collection using the conversational search history. Currently proposed models for these datasets incorporate history in retrieval by appending the last N turns to the current question before encoding. We propose to use another history selection approach that dynamically selects and weighs history turns using the attention mechanism for question embedding. The novelty of our approach lies in experimenting with soft attention-based history selection approach in an open-retrieval setting.

cs.IR

Coswara -- A Database of Breathing, Cough, and Voice Sounds for COVID-19 Diagnosis

The COVID-19 pandemic presents global challenges transcending boundaries of country, race, religion, and economy. The current gold standard method for COVID-19 detection is the reverse transcription polymerase chain reaction (RT-PCR) testing. However, this method is expensive, time-consuming, and violates social distancing. Also, as the pandemic is expected to stay for a while, there is a need for an alternate diagnosis tool which overcomes these limitations, and is deployable at a large scale. The prominent symptoms of COVID-19 include cough and breathing difficulties. We foresee that respiratory sounds, when analyzed using machine learning techniques, can provide useful insights, enabling the design of a diagnostic tool. Towards this, the paper presents an early effort in creating (and analyzing) a database, called Coswara, of respiratory sounds, namely, cough, breath, and voice. The sound samples are collected via worldwide crowdsourcing using a website application. The curated dataset is released as open access. As the pandemic is evolving, the data collection and analysis is a work in progress. We believe that insights from analysis of Coswara can be effective in enabling sound based technology solutions for point-of-care diagnosis of respiratory infection, and in the near future this can help to diagnose COVID-19.

eess.AS

Efficient cache oblivious algorithms for randomized divide-and-conquer on the multicore model

In this paper we present randomized algorithms for sorting and convex hull that achieves optimal performance (for speed-up and cache misses) on the multicore model with private cache model. Our algorithms are cache oblivious and generalize the randomized divide and conquer strategy given by Reischuk and Reif and Sen. Although the approach yielded optimal speed-up in the PRAM model, we require additional techniques to optimize cache-misses in an oblivious setting. Under a mild assumption on input and number of processors our algorithm will have optimal time and cache misses with high probability. Although similar results have been obtained recently for sorting, we feel that our approach is simpler and general and we apply it to obtain an optimal parallel algorithm for 3D convex hulls with similar bounds. We also present a simple randomized processor allocation technique without the explicit knowledge of the number of processors that is likely to find additional applications in resource oblivious environments.

cs.DS