Searcharxiv⌕ Search

arXiv subjects

Rahul Sharma

Publications and source records attributed to Rahul Sharma.

At least 73 records · Page 4Linked to original sources

AstroSat and NuSTAR observations of XTE J1739-285 during the 2019-2020 outburst

We report results from a study of XTE J1739-285, a transient neutron star low mass X-ray binary observed with AstroSat and NuSTAR during its 2019-2020 outburst. We detected accretion-powered X-ray pulsations at 386 Hz during very short intervals (0.5--1 s) of X-ray flares. These flares were observed during the 2019 observation of XTE J1739-285. During this observation, we also observed a correlation between intensity and hardness ratios, suggesting an increase in hardness with the increase in intensity. Moreover, a thermonuclear X-ray burst detected in our AstroSat observation during the 2020 outburst revealed the presence of coherent burst oscillations at 383 Hz during its decay phase. The frequency drift of 3 Hz during X-ray burst can be explained with r modes. Thus, making XTE J1739-285 belong to a subset of NS-LMXBs which exhibit both nuclear- and accretion-powered pulsations. The power density spectrum created using the AstroSat-LAXPC observations in 2020 showed the presence of a quasi-periodic oscillation at ~ 0.83 Hz. Our X-ray spectroscopy revealed significant changes in the spectra during the 2019 and 2020 outburst. We found a broad iron line emission feature in the X-ray spectrum during the 2020 observation, while this feature was relatively narrow and has a lower equivalent width in 2019,~when the source was accreting at higher rates than 2020.

astro-ph.HE↗

Machine Learning for Optical Motion Capture-driven Musculoskeletal Modelling from Inertial Motion Capture Data

Marker-based Optical Motion Capture (OMC) systems and associated musculoskeletal (MSK) modelling predictions offer non-invasively obtainable insights into in vivo joint and muscle loading, aiding clinical decision-making. However, an OMC system is lab-based, expensive, and requires a line of sight. Inertial Motion Capture (IMC) systems are widely-used alternatives, which are portable, user-friendly, and relatively low-cost, although with lesser accuracy. Irrespective of the choice of motion capture technique, one needs to use an MSK model to obtain the kinematic and kinetic outputs, which is a computationally expensive tool increasingly well approximated by machine learning (ML) methods. Here, we present an ML approach to map experimentally recorded IMC data to the human upper-extremity MSK model outputs computed from ('gold standard') OMC input data. Essentially, we aim to predict higher-quality MSK outputs from the much easier-to-obtain IMC data. We use OMC and IMC data simultaneously collected for the same subjects to train different ML architectures that predict OMC-driven MSK outputs from IMC measurements. In particular, we employed various neural network (NN) architectures, such as Feed-Forward Neural Networks (FFNNs) and Recurrent Neural Networks (RNNs) (vanilla, Long Short-Term Memory, and Gated Recurrent Unit) and searched for the best-fit model through an exhaustive search in the hyperparameters space in both subject-exposed (SE) & subject-naive (SN) settings. We observed a comparable performance for both FFNN & RNN models, which have a high degree of agreement (ravg, SE, FFNN = 0.90+/-0.19, ravg, SE, RNN = 0.89+/-0.17, ravg, SN, FFNN = 0.84+/-0.23, & ravg, SN, RNN = 0.78+/-0.23) with the desired OMC-driven MSK estimates for held-out test data. Mapping IMC inputs to OMC-driven MSK outputs using ML models could be instrumental in transitioning MSK modelling from 'lab to field'.

cs.LG↗

Machine learning techniques for the Schizophrenia diagnosis: A comprehensive review and future research directions

Schizophrenia (SCZ) is a brain disorder where different people experience different symptoms, such as hallucination, delusion, flat-talk, disorganized thinking, etc. In the long term, this can cause severe effects and diminish life expectancy by more than ten years. Therefore, early and accurate diagnosis of SCZ is prevalent, and modalities like structural magnetic resonance imaging (sMRI), functional MRI (fMRI), diffusion tensor imaging (DTI), and electroencephalogram (EEG) assist in witnessing the brain abnormalities of the patients. Moreover, for accurate diagnosis of SCZ, researchers have used machine learning (ML) algorithms for the past decade to distinguish the brain patterns of healthy and SCZ brains using MRI and fMRI images. This paper seeks to acquaint SCZ researchers with ML and to discuss its recent applications to the field of SCZ study. This paper comprehensively reviews state-of-the-art techniques such as ML classifiers, artificial neural network (ANN), deep learning (DL) models, methodological fundamentals, and applications with previous studies. The motivation of this paper is to benefit from finding the research gaps that may lead to the development of a new model for accurate SCZ diagnosis. The paper concludes with the research finding, followed by the future scope that directly contributes to new research directions.

cs.LG↗

The AstroSat observation of accreting millisecond X-ray pulsar SAX J1808.4-3658 during its 2019 outburst

We report on the analysis of the AstroSat dataset of the accreting millisecond X-ray pulsar SAX J1808.4-3658, obtained during its 2019 outburst. We found coherent pulsations at $\sim 401$ Hz and an orbital solution consistent with previous studies. The 3-20 keV pulse profile can be well fitted with three harmonically related sinusoidal components with background-corrected fractional amplitude of $\sim 3.5 \%$, $\sim 1.2 \%$ and $\sim 0.37 \%$ for fundamental, second and third harmonic, respectively. Our energy-resolved pulse profile evolution study indicate a strong energy dependence. We also observed a soft lag in fundamental and hard lag during its harmonic. The broadband spectrum of SAX J1808.4-3658 can be well described with a combination of thermal emission component with $kT \sim 1$ keV, a thermal Comptonization ($Γ\sim 1.67$) from the hot corona and broad emission lines due to Fe.

astro-ph.HE↗

Broadband mHz QPOs and spectral study of LMC X$-$4 with AstroSat

We report the results of broadband timing and spectral analysis of data from an AstroSat observation of the High Mass X-ray binary LMC X$-$4. The Large Area X-ray Proportional Counter (LAXPC) and Soft X-ray Telescope (SXT) instruments on-board the AstroSat observed the source in August 2016. A complete X-ray eclipse was detected with the LAXPC. The 3$-$40 keV power density spectrum showed the presence of coherent pulsations along with a $\sim 26$ mHz quasi-periodic oscillation feature. The spectral properties of LMC X$-$4 were derived from a joint analysis of the SXT and LAXPC spectral data. The 0.5$-$25 keV persistent spectrum comprised of an absorbed high energy cutoff power law with photon index of $Γ\sim$ 0.8 and cutoff at $\sim$16 keV, a soft thermal component with kT$_{BB} \sim$ 0.14 keV and Gaussian components corresponding to Fe K$_α$, Ne \textsc{ix} and Ne \textsc{x} emission lines. Assuming a source distance of 50 kpc, we determined 0.5--25 keV luminosity to be $\sim 2 \times 10^{38}$ erg s$^{-1}$.

astro-ph.HE↗

Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection

Active speaker detection in videos addresses associating a source face, visible in the video frames, with the underlying speech in the audio modality. The two primary sources of information to derive such a speech-face relationship are i) visual activity and its interaction with the speech signal and ii) co-occurrences of speakers' identities across modalities in the form of face and speech. The two approaches have their limitations: the audio-visual activity models get confused with other frequently occurring vocal activities, such as laughing and chewing, while the speakers' identity-based methods are limited to videos having enough disambiguating information to establish a speech-face association. Since the two approaches are independent, we investigate their complementary nature in this work. We propose a novel unsupervised framework to guide the speakers' cross-modal identity association with the audio-visual activity for active speaker detection. Through experiments on entertainment media videos from two benchmark datasets, the AVA active speaker (movies) and Visual Person Clustering Dataset (TV shows), we show that a simple late fusion of the two approaches enhances the active speaker detection performance.

cs.MM↗

MinUn: Accurate ML Inference on Microcontrollers

Running machine learning inference on tiny devices, known as TinyML, is an emerging research area. This task requires generating inference code that uses memory frugally, a task that standard ML frameworks are ill-suited for. A deployment framework for TinyML must be a) parametric in the number representation to take advantage of the emerging representations like posits, b) carefully assign high-precision to a few tensors so that most tensors can be kept in low-precision while still maintaining model accuracy, and c) avoid memory fragmentation. We describe MinUn, the first TinyML framework that holistically addresses these issues to generate efficient code for ARM microcontrollers (e.g., Arduino Uno, Due and STM32H747) that outperforms the prior TinyML frameworks.

cs.LG↗

Eclipse Timings of the LMXB XTE J1710-281 : Discovery of a third orbital period glitch

We present an updated measurement of orbital period evolution of LMXB XTE J1710-281 by using eclipse timing technique. Using data obtained with XMM-Newton, Suzaku, RXTE, Chandra and AstroSat observatories, we report 21 new measurements of X-ray mid-eclipse times. We have discovered a third orbital period glitch in XTE J1710-281 with an F-test false alarm probability of ~0.7% for occurrence of the third glitch and report detection of four distinct epochs of orbital period in this system. This work presents a more robust estimation of occurrence of the second orbital period glitch. However, the epoch of occurrence of the third glitch is poorly constrained, between MJD 55726 to 56402. We have put lower limits of 1.48 ms, 0.97 ms and 0.45 ms, on sudden changes in orbital period between the successive epochs. We discuss the implications of our findings in context of magnetic nature of the companion star and possible scattering events with circum-binary objects around this binary system.

astro-ph.HE↗

Unsupervised active speaker detection in media content using cross-modal information

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial images. We leverage speaker identity information from speech and faces, and formulate active speaker detection as a speech-face assignment task such that the active speaker's face and the underlying speech identify the same person (character). We express the speech segments in terms of their associated speaker identity distances, from all other speech segments, to capture a relative identity structure for the video. Then we assign an active speaker's face to each speech segment from the concurrently appearing faces such that the obtained set of active speaker faces displays a similar relative identity structure. Furthermore, we propose a simple and effective approach to address speech segments where speakers are present off-screen. We evaluate the proposed system on three benchmark datasets -- Visual Person Clustering dataset, AVA-active speaker dataset, and Columbia dataset -- consisting of videos from entertainment and broadcast media, and show competitive performance to state-of-the-art fully supervised methods.

eess.IV↗

Efficient ML Models for Practical Secure Inference

ML-as-a-service continues to grow, and so does the need for very strong privacy guarantees. Secure inference has emerged as a potential solution, wherein cryptographic primitives allow inference without revealing users' inputs to a model provider or model's weights to a user. For instance, the model provider could be a diagnostics company that has trained a state-of-the-art DenseNet-121 model for interpreting a chest X-ray and the user could be a patient at a hospital. While secure inference is in principle feasible for this setting, there are no existing techniques that make it practical at scale. The CrypTFlow2 framework provides a potential solution with its ability to automatically and correctly translate clear-text inference to secure inference for arbitrary models. However, the resultant secure inference from CrypTFlow2 is impractically expensive: Almost 3TB of communication is required to interpret a single X-ray on DenseNet-121. In this paper, we address this outstanding challenge of inefficiency of secure inference with three contributions. First, we show that the primary bottlenecks in secure inference are large linear layers which can be optimized with the choice of network backbone and the use of operators developed for efficient clear-text inference. This finding and emphasis deviates from many recent works which focus on optimizing non-linear activation layers when performing secure inference of smaller networks. Second, based on analysis of a bottle-necked convolution layer, we design a X-operator which is a more efficient drop-in replacement. Third, we show that the fast Winograd convolution algorithm further improves efficiency of secure inference. In combination, these three optimizations prove to be highly effective for the problem of X-ray interpretation trained on the CheXpert dataset.

cs.CR↗

Change in spin-down rate and detection of emission line in HMXB 4U 2206+54 with AstroSat observation

This work presents timing and spectral analysis of 4U 2206+54 using data obtained from LAXPC instrument onboard India's AstroSat mission. This source was observed with AstroSat in September 2016 and October 2016. We report detection of $5648 (4)$ s pulsations at MJD 57669 in the latter observation of 4U 2206+54. The pulse profile is sinusoidal and the inherent shape is independent of energy up to 30 keV. The pulse fraction increases with energy from $\sim 0.5$% to $\sim 0.8$%. We report an updated spin down rate of $2.95 (14) \times 10^{-7}$ s s$^{-1}$. This is about 0.40 times smaller than the previously reported long term value. The energy spectrum is best modelled with an absorbed power-law with high energy exponential cut-off. We have detected presence of broad emission line in 4U 2206+54 at an energy of $7$ keV with equivalent width of $\sim 0.4$ keV.

astro-ph.HE↗

Small-scale solar jet formation and their associated waves and instabilities

Studies on small-scale jets' formation, propagation, evolution, and role, such as type I and II spicules, mottles, and fibrils in the lower solar atmosphere's energetic balance, have progressed tremendously thanks to the combination of detailed observations and sophisticated mathematical modelling. This review provides a survey of the current understanding of jets, their formation in the solar lower atmosphere, and their evolution from observational, numerical, and theoretical perspectives. First, we review some results to describe the jet properties, acquired numerically, analytically and through high-spatial and temporal resolution observations. Further on, we discuss the role of hydrodynamic and magnetohydrodynamic instabilities, namely Rayleigh-Taylor and Kelvin-Helmholtz instabilities, in jet evolution and their role in the energy transport through the solar atmosphere in fully and partially ionised plasmas. Finally, we discuss several mechanisms of magnetohydrodynamic wave generation, propagation, and energy transport in the context of small-scale solar jets in detail. This review identifies several gaps in the understanding of small-scale solar jets and some misalignments between the observational studies and knowledge acquired through theoretical studies and numerical modelling. It is to be expected that these gaps will be closed with the advent of high-resolution observational instruments, such as Daniel K. Inouye Solar Telescope, Solar Orbiter, Parker Solar Probe, and Solar CubeSats for Linked Imaging Spectropolarimetry, combined with further theoretical and computational developments.

astro-ph.SR↗

Federated Learning with Noisy User Feedback

Machine Learning (ML) systems are getting increasingly popular, and drive more and more applications and services in our daily life. This has led to growing concerns over user privacy, since human interaction data typically needs to be transmitted to the cloud in order to train and improve such systems. Federated learning (FL) has recently emerged as a method for training ML models on edge devices using sensitive user data and is seen as a way to mitigate concerns over data privacy. However, since ML models are most commonly trained with label supervision, we need a way to extract labels on edge to make FL viable. In this work, we propose a strategy for training FL models using positive and negative user feedback. We also design a novel framework to study different noise patterns in user feedback, and explore how well standard noise-robust objectives can help mitigate this noise when training models in a federated setting. We evaluate our proposed training setup through detailed experiments on two text classification datasets and analyze the effects of varying levels of user reliability and feedback noise on model performance. We show that our method improves substantially over a self-training baseline, achieving performance closer to models trained with full supervision.

cs.LG↗

On the Existence of Photoluminescence and Room-Temperature Spin Polarization in Ambipolar V doped MoS$_2$ Monolayers

Opto-spintronics is an emerging field where ultra-thin magnetic-semiconductors having high spin-valley coupling play an important role. Here, we demonstrate substitutional vanadium (V) doping in MoS$_2$ lattice in different extent, leading to the coexistence of photoluminescence (PL), valleypolarization (~32%), and valley splitting (~28 meV shift in PL with helicity $σ^+$ and $σ^-$ of light excitation). A large V doping causes semiconductor to metal transition in MoS$_2$ but with medium level causing the existence of photoluminescence with high spin polarization. The ambipolar nature of medium level V doped MoS$_2$ is shown here indicating its potential as an opto-electronic material. The presence of V-dopants and their different level of content are proven by both spectroscopic and microscopic methods.A detailed temperature and power dependent photoluminescence studies along with density functional theory-based calculations in support unravels the emergence of the co-existence of spin-valley coupling and photoluminescence. This study shows the potential of doping MoS$_2$ for deriving new materials for next generation room temperature opto-spintronics.

physics.app-ph↗

Using Active Speaker Faces for Diarization in TV shows

Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understanding. Automated audio-based speaker diarization of entertainment media poses challenges due to the diverse acoustic conditions present in media content, be it background music, overlapping speakers, or sound effects. At the same time, speaking faces in the visual modality provide complementary information and not prone to the errors seen in the audio modality. In this paper, we address the problem of speaker diarization in TV shows using the active speaker faces. We perform face clustering on the active speaker faces and show superior speaker diarization performance compared to the state-of-the-art audio-based diarization methods. We additionally report a systematic analysis of the impact of active speaker face detection quality on the diarization performance. We also observe that a moderately well-performing active speaker system could outperform the audio-based diarization systems.

cs.MM↗

Audio visual character profiles for detecting background characters in entertainment media

An essential goal of computational media intelligence is to support understanding how media stories -- be it news, commercial or entertainment media -- represent and reflect society and these portrayals are perceived. People are a central element of media stories. This paper focuses on understanding the representation and depiction of background characters in media depictions, primarily movies and TV shows. We define the background characters as those who do not participate vocally in any scene throughout the movie and address the problem of localizing background characters in videos. We use an active speaker localization system to extract high-confidence face-speech associations and generate audio-visual profiles for talking characters in a movie by automatically clustering them. Using a face verification system, we then prune all the face-tracks which match any of the generated character profiles and obtain the background character face-tracks. We curate a background character dataset which provides annotations for background character for a set of TV shows, and use it to evaluate the performance of the background character detection framework.

cs.CV↗

Neutral atom scattering based mapping of atomically thin layers

Imaging surfaces using low energy neutral atom scattering is a relatively recent development in the field of microscopy. In this work we demonstrate that this technique is sensitive enough to distinguish films as thin as a single monolayer from the underlying substrate. Using collimated beams of He and Kr atoms as an incident probe on MoS$_2$ films grown on SiO$_2$/Si substrate, we observe systematic changes in the scattered atom flux which allows us to map the thin MoS$_2$ films. Measurements carried out by varying incidence energy using both He and Kr provides insights into the details of atom-surface collision dynamics and its role in contrast generation.

physics.chem-ph↗

Jigsaw: Large Language Models meet Program Synthesis

Large pre-trained language models such as GPT-3, Codex, and Google's language model are now capable of generating code from natural language specifications of programmer intent. We view these developments with a mixture of optimism and caution. On the optimistic side, such large language models have the potential to improve productivity by providing an automated AI pair programmer for every programmer in the world. On the cautionary side, since these large language models do not understand program semantics, they offer no guarantees about quality of the suggested code. In this paper, we present an approach to augment these large language models with post-processing steps based on program analysis and synthesis techniques, that understand the syntax and semantics of programs. Further, we show that such techniques can make use of user feedback and improve with usage. We present our experiences from building and evaluating such a tool jigsaw, targeted at synthesizing code for using Python Pandas API using multi-modal inputs. Our experience suggests that as these large language models evolve for synthesizing code from intent, jigsaw has an important role to play in improving the accuracy of the systems.

cs.SE↗