Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 127 records · Page 7Linked to original sources

Multi-wavelength study of the luminous GRB 210619B observed with Fermi and ASIM

We report on detailed multi-wavelength observations and analysis of the very bright and long GRB 210619B, detected by the Atmosphere-Space Interactions Monitor (ASIM) installed on the International Space Station (ISS) and the Gamma-ray Burst Monitor (GBM) on-board the Fermi mission. Our main goal is to understand the radiation mechanisms and jet composition of GRB 210619B. With a measured redshift of $z$ = 1.937, we find that GRB 210619B falls within the 10 most luminous bursts observed by Fermi so far. The energy-resolved prompt emission light curve of GRB 210619B exhibits an extremely bright hard emission pulse followed by softer/longer emission pulses. The low-energy photon indices ($α_{\rm pt}$) values obtained using the time-resolved spectral analysis of the burst suggest a transition between the thermal (during harder pulse) to non-thermal (during softer pulse) outflow. We examine the correlation between spectral parameters and find that both peak energy and $α_{\rm pt}$ exhibit the flux tracking pattern. The late time broadband photometric dataset can be explained within the framework of the external forward shock model with $ν_m$ $< ν_c$ $< ν_{x}$ (where $ν_m$, $ν_c$, and $ν_{x}$ are the synchrotron peak, cooling-break, and X-ray frequencies, respectively) spectral regime supporting a rarely observed hard electron energy index ($p<$ 2). We find moderate values of host extinction of E(B-V) = 0.14 $\pm$ 0.01 mag for the Small Magellanic Cloud (SMC) extinction law. In addition, we also report late-time optical observations with the 10.4 m GTC placing deep upper limits for the host galaxy ($z$=1.937), favouring a faint, dwarf host for the burst.

astro-ph.HE↗

Prompt emission and early optical afterglow of VHE detected GRB 201015A and GRB 201216C: onset of the external forward shock

We present a detailed prompt emission and early optical afterglow analysis of the two very high energy (VHE) detected bursts GRB 201015A and GRB 201216C, and their comparison with a subset of similar bursts. Time-resolved spectral analysis of multi-structured GRB 201216C using the Bayesian binning algorithm revealed that during the entire duration of the burst, the low energy spectral index ($α_{\rm pt}$) remained below the limit of the synchrotron line of death. However, statistically some of the bins supported the additional thermal component. Additionally, the evolution of spectral parameters showed that both peak energy (Ep) and $α_{\rm pt}$ tracked the flux. These results were further strengthened using the values of the physical parameters obtained by synchrotron modeling of the data. Our earliest optical observations of both bursts using FRAM-ORM and BOOTES robotic telescopes displayed a smooth bump in their early optical light curves, consistent with the onset of the afterglow due to synchrotron emission from an external forward shock. Using the observed optical peak, we constrained the initial bulk Lorentz factors of GRB 201015A and GRB 201216C to $Γ_0$ = 204 and $Γ_0$ = 310, respectively. The present early optical observations are the earliest known observations constraining outflow parameters and our analysis indicate that VHE-detected bursts could have a diverse range of observed luminosity within the detectable redshift range of present VHE facilities.

astro-ph.HE↗

Is the Elephant Flying? Resolving Ambiguities in Text-to-Image Generative Models

Natural language often contains ambiguities that can lead to misinterpretation and miscommunication. While humans can handle ambiguities effectively by asking clarifying questions and/or relying on contextual cues and common-sense knowledge, resolving ambiguities can be notoriously hard for machines. In this work, we study ambiguities that arise in text-to-image generative models. We curate a benchmark dataset covering different types of ambiguities that occur in these systems. We then propose a framework to mitigate ambiguities in the prompts given to the systems by soliciting clarifications from the user. Through automatic and human evaluations, we show the effectiveness of our framework in generating more faithful images aligned with human intention in the presence of ambiguities.

cs.CL↗

Photometric studies on the host galaxies of gamma-ray bursts using 3.6m Devasthal Optical Telescope

In this article, we present multi-band photometric observations and analysis of the host galaxies for a sample of five interesting gamma-ray bursts (GRBs) observed using the 3.6m Devasthal Optical Telescope (DOT) and the back-end instruments. The host galaxy observations of GRBs provide unique opportunities to estimate the stellar mass, ages, star-formation rates, and other vital properties of the burst environments and hence progenitors. We performed a detailed spectral energy distribution (SED) modeling of the five host galaxies using an advanced tool called Prospector, a stellar population synthesis model. Furthermore, we compared the results with a larger sample of well-studied host galaxies of GRBs, supernovae, and normal star-forming galaxies. Our SED modeling suggests that GRB 130603B, GRB 140102A, GRB 190829A, and GRB 200826A have massive host galaxies with high star formation rates (SFRs). On the other hand, a supernovae-connected GRB 030329 has a rare low-mass galaxy with a low star formation rate. We also find that GRB 190829A has the highest (in our sample) amount of visual dust extinction and gas in its local environment of the host, suggesting that the observed very high energy emission from this burst might have a unique local environment. Broadly, the five GRBs in our sample satisfy the typical correlations between host galaxies parameters and these physical parameters are more common to normal star-forming galaxies at the high-redshift Universe. Our results also demonstrate the capabilities of 3.6m DOT and the back-end instruments for the deeper photometric studies of the host galaxies of energetic transients such as GRBs, supernovae, and other transients in the long run.

astro-ph.HE↗

Design Considerations For Hypothesis Rejection Modules In Spoken Language Understanding Systems

Spoken Language Understanding (SLU) systems typically consist of a set of machine learning models that operate in conjunction to produce an SLU hypothesis. The generated hypothesis is then sent to downstream components for further action. However, it is desirable to discard an incorrect hypothesis before sending it downstream. In this work, we present two designs for SLU hypothesis rejection modules: (i) scheme R1 that performs rejection on domain specific SLU hypothesis and, (ii) scheme R2 that performs rejection on hypothesis generated from the overall SLU system. Hypothesis rejection modules in both schemes reject/accept a hypothesis based on features drawn from the utterance directed to the SLU system, the associated SLU hypothesis and SLU confidence score. Our experiments suggest that both the schemes yield similar results (scheme R1: 2.5% FRR @ 4.5% FAR, scheme R2: 2.5% FRR @ 4.6% FAR), with the best performing systems using all the available features. We argue that while either of the rejection schemes can be chosen over the other, they carry some inherent differences which need to be considered while making this choice. Additionally, we incorporate ASR features in the rejection module (obtaining an 1.9% FRR @ 3.8% FAR) and analyze the improvements.

cs.CL↗

An Analysis of the Effects of Decoding Algorithms on Fairness in Open-Ended Language Generation

Several prior works have shown that language models (LMs) can generate text containing harmful social biases and stereotypes. While decoding algorithms play a central role in determining properties of LM generated text, their impact on the fairness of the generations has not been studied. We present a systematic analysis of the impact of decoding algorithms on LM fairness, and analyze the trade-off between fairness, diversity and quality. Our experiments with top-$p$, top-$k$ and temperature decoding algorithms, in open-ended language generation, show that fairness across demographic groups changes significantly with change in decoding algorithm's hyper-parameters. Notably, decoding algorithms that output more diverse text also output more texts with negative sentiment and regard. We present several findings and provide recommendations on standardized reporting of decoding details in fairness evaluations and optimization of decoding algorithms for fairness alongside quality and diversity.

cs.CL↗

Differentially Private Decoding in Large Language Models

Recent large-scale natural language processing (NLP) systems use a pre-trained Large Language Model (LLM) on massive and diverse corpora as a headstart. In practice, the pre-trained model is adapted to a wide array of tasks via fine-tuning on task-specific datasets. LLMs, while effective, have been shown to memorize instances of training data thereby potentially revealing private information processed during pre-training. The potential leakage might further propagate to the downstream tasks for which LLMs are fine-tuned. On the other hand, privacy-preserving algorithms usually involve retraining from scratch, which is prohibitively expensive for LLMs. In this work, we propose a simple, easy to interpret, and computationally lightweight perturbation mechanism to be applied to an already trained model at the decoding stage. Our perturbation mechanism is model-agnostic and can be used in conjunction with any LLM. We provide theoretical analysis showing that the proposed mechanism is differentially private, and experimental results showing a privacy-utility trade-off.

cs.CL↗

SN 2016iyc: A Type IIb supernova arising from a low-mass progenitor

In this work, photometric and spectroscopic analyses of a very low-luminosity Type IIb supernova (SN) 2016iyc have been performed. SN 2016iyc lies near the faint end among the distribution of similar supernovae (SNe). Given lower ejecta mass ($M_{\rm ej}$) and low nickel mass ($M_{\rm Ni}$) from the literature, combined with SN 2016iyc lying near the faint end, one-dimensional stellar evolution models of 9 - 14 M$_{\odot}$ zero-age main-sequence (ZAMS) stars as the possible progenitors of SN 2016iyc have been performed using the publicly available code MESA. Moreover, synthetic explosions of the progenitor models have been simulated using the hydrodynamic evolution codes STELLA and SNEC. The bolometric luminosity light curve and photospheric velocities produced through synthetic explosions of ZAMS stars of mass in the range 12 - 13 M$_{\odot}$ having a pre-supernova radius $R_{\mathrm{0}} =$ (240 - 300) R$_{\odot}$, with $M_{\rm ej} =$ (1.89 - 1.93) M$_{\odot}$, explosion energy $E_{\rm exp} = $ (0.28 - 0.35) $\times 10^{51}$ erg, and $M_{\rm Ni} < 0.09$ M$_{\odot}$, are in good agreement with observations; thus, SN 2016iyc probably exploded from a progenitor near the lower mass limits for SNe IIb. Finally, hydrodynamic simulations of the explosions of SN 2016gkg and SN 2011fu have also been performed to compare intermediate- and high-luminosity examples among well-studied SNe IIb. The results of progenitor modelling and synthetic explosions for SN 2016iyc, SN 2016gkg, and SN 2011fu exhibit a diverse range of mass for the possible progenitors of SNe IIb.

astro-ph.HE↗

AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model

In this work, we demonstrate that multilingual large-scale sequence-to-sequence (seq2seq) models, pre-trained on a mixture of denoising and Causal Language Modeling (CLM) tasks, are more efficient few-shot learners than decoder-only models on various tasks. In particular, we train a 20 billion parameter multilingual seq2seq model called Alexa Teacher Model (AlexaTM 20B) and show that it achieves state-of-the-art (SOTA) performance on 1-shot summarization tasks, outperforming a much larger 540B PaLM decoder model. AlexaTM 20B also achieves SOTA in 1-shot machine translation, especially for low-resource languages, across almost all language pairs supported by the model (Arabic, English, French, German, Hindi, Italian, Japanese, Marathi, Portuguese, Spanish, Tamil, and Telugu) on Flores-101 dataset. We also show in zero-shot setting, AlexaTM 20B outperforms GPT3 (175B) on SuperGLUE and SQuADv2 datasets and provides SOTA performance on multilingual tasks such as XNLI, XCOPA, Paws-X, and XWinograd. Overall, our results present a compelling case for seq2seq models as a powerful alternative to decoder-only models for Large-scale Language Model (LLM) training.

cs.CL↗

A decomposition theorem for 0-cycles and applications

We prove a decomposition theorem for the cohomological Chow group of 0-cycles on the double of a quasi-projective $R_1$-scheme over a field along a closed subscheme, in terms of the Chow groups, with and without modulus, of the scheme. This yields a significant generalization of the decomposition theorem of Binda-Krishna. As applications, we prove a moving lemma for Chow groups with modulus and an analogue of Bloch's formula for 0-cycles with modulus on singular surfaces. The latter extends a previous result of Binda-Krishna-Saito.

math.AG↗

Hard X-ray polarization catalog for a 5-year sample of Gamma-Ray Bursts using AstroSat CZT-Imager

Cadmium Zinc Telluride Imager (CZTI) aboard AstroSat has been regularly detecting Gamma-Ray Bursts (GRBs) since its launch in 2015. Its sensitivity to polarization measurements at energies above 100 keV allows CZTI to attempt spectro-polarimetric studies of GRBs. Here, we present the first catalog of GRB polarization measurements made by CZTI during its first five years of operation. This presents the time integrated polarization measurements of the prompt emission of 20 GRBs in the energy range 100-600 keV. The sample includes the bright GRBs which were detected within an angle range of 0-60 degree and 120-180 degree where the instrument has useful polarization sensitivity and is less prone to systematics. We implement a few new modifications in the analysis to enhance polarimetric sensitivity of the instrument. Majority of the GRBs in the sample are found to possess less / null polarization across the total bursts' duration in contrast to a small fraction of five GRBs exhibiting high polarization. The low polarization across the bursts can be speculated to be either due to the burst being intrinsically weakly polarized or due to varying polarization angle within the burst even when it is highly polarized. In comparison to POLAR measurements, CZTI has detected a larger number of cases with high polarization. This may be a consequence of the higher energy window of CZTI observations which results in the sampling of smaller duration of burst emissions in contrast to POLAR, thereby, probing emissions of less temporal variations of polarization properties.

astro-ph.HE↗

Tale of GRB 171010A/SN 2017htp and GRB 171205A/SN 2017iuk: Magnetar origin?

We present late-time optical follow-up observations of GRB 171010A/SN 2017htp ($z$ = 0.33) and low-luminosity GRB 171205A/SN 2017iuk ($z$ = 0.037) acquired using the 4K$\times$4K CCD Imager mounted at the 3.6m Devasthal Optical Telescope (3.6m DOT) along with the prompt emission data analysis of these two interesting bursts. The prompt characteristics (other than brightness) such as spectral hardness, T$_{90}$, and minimum variability time-scale are comparable for both the bursts. The isotropic $X$-ray and kinetic energies of the plateau phase of GRB 171205A are found to be less than the maximum energy budget of magnetars, supporting magnetar as a central engine powering source. The new optical data of SN 2017htp and SN 2017iuk presented here, along with published ones, indicate that SN 2017htp is one of the brightest and SN 21017iuk is among the faintest GRB associated SNe (GRB-SNe). Semi-analytical light-curve modelling of SN 2017htp, SN 2017iuk and only known GRB associated superluminous supernova (SLSN 2011kl) are performed using the $\texttt{MINIM}$ code. The model with a spin-down millisecond magnetar as a central engine powering source nicely reproduced the bolometric light curves of all three GRB-SNe mentioned above. The magnetar central engines for SN 2017htp, SN 2017iuk, and SLSN 2011kl exhibit values of initial spin periods higher and magnetic fields closer to those observed for long GRBs and H-deficient SLSNe. Detection of these rare events at such late epochs also demonstrates the capabilities of the 3.6m DOT for deep imaging considering longitudinal advantage in the era of time-domain astronomy.

astro-ph.HE↗

Reciprocity for Kato-Saito idele class group with modulus

We introduce an etale fundamental group with modulus and construct a reciprocity homomorphism from the Kato-Saito idele class group with modulus to this fundamental group. This is the K-theoretic analogue of the reciprocity for the cycle-theoretic idele class group with modulus due to Kerz-Saito, and plays a central role in showing the isomorphism between the two idele class groups. It also provides a new interpretation of the already known etale fundamental group with modulus due to Deligne and Laumon.

math.AG↗

Core-collapse supernova from a possible progenitor star of 100 M$_{\odot}$

In this work, we study the synthetic explosions of a massive star. We take a 100 M$_{\odot}$ zero--age main--sequence (ZAMS) star and evolve it until the onset of core-collapse using {\tt MESA}. Then, the resulting star model is exploded using the publicly available stellar explosion code, {\tt STELLA}. The outputs of {\tt STELLA} calculations provide us the bolometric light curve and photospheric velocity evolution along with other physical properties of the underlying supernova. In this paper, the effects of having large Hydrogen-envelope on the supernova light curve have been explored. We also explore the effects of the presence of different amounts of nickel mass and the effect of changing the explosion energy of the resulting supernovae from such heavy progenitors, on their bolometric light curves and photospheric velocities.

astro-ph.HE↗

Analyses of Hydrogen-stripped core-collapse supernovae using MOSFiT and MESA based tools

In this work, we employ two publicly available analysis tools to study four hydrogen(H)--stripped core--collapse supernovae (CCSNe) namely, SN 2009jf, iPTF13bvn, SN 2015ap, and SN 2016bau. We use the Modular Open-Source Fitter for Transients ({\tt MOSFiT}) to model the multi band light curves. {\tt MOSFiT} analyses show ejecta masses (log M$_{ej}$) of $0.80_{-0.13}^{+0.18}$ M$_{\odot}$, $0.15_{-0.09}^{+0.13}$ M$_{\odot}$, $0.19_{-0.03}^{+0.03}$ M$_{\odot}$, and $0.19_{+0.02}^{-0.01}$ M$_{\odot}$ for SN 2009jf, iPTF13vn, SN 2015ap, and SN 2016au, respectively. Later, Modules for Experiments in Stellar Astrophysics ({\tt MESA}), is used to construct models of stars from pre-main sequence upto core collapse which serve as the possible progenitors of these H-stripped CCSNe. Based on literature, we model a 12 M$_{\odot}$ ZAMS star as the possible progenitor for iPTF13vn, SN 2015ap, and SN 2016bau while a 20 M$_{\odot}$ ZAMS star is modeled as the possible progenitor for SN 2009jf. Glimpses of stellar engineering and the physical properties of models at various stages of their lifetime have been presented to demonstrate the usefulness of these analysis threads to understand the observed properties of several classes of transients in detail.

astro-ph.HE↗

FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks

Increasing concerns and regulations about data privacy and sparsity necessitate the study of privacy-preserving, decentralized learning methods for natural language processing (NLP) tasks. Federated learning (FL) provides promising approaches for a large number of clients (e.g., personal devices or organizations) to collaboratively learn a shared global model to benefit all clients while allowing users to keep their data locally. Despite interest in studying FL methods for NLP tasks, a systematic comparison and analysis is lacking in the literature. Herein, we present the FedNLP, a benchmarking framework for evaluating federated learning methods on four different task formulations: text classification, sequence tagging, question answering, and seq2seq. We propose a universal interface between Transformer-based language models (e.g., BERT, BART) and FL methods (e.g., FedAvg, FedOPT, etc.) under various non-IID partitioning strategies. Our extensive experiments with FedNLP provide empirical comparisons between FL methods and helps us better understand the inherent challenges of this direction. The comprehensive analysis points to intriguing and exciting future research aimed at developing FL methods for NLP tasks.

cs.CL↗

Federated Learning with Noisy User Feedback

Machine Learning (ML) systems are getting increasingly popular, and drive more and more applications and services in our daily life. This has led to growing concerns over user privacy, since human interaction data typically needs to be transmitted to the cloud in order to train and improve such systems. Federated learning (FL) has recently emerged as a method for training ML models on edge devices using sensitive user data and is seen as a way to mitigate concerns over data privacy. However, since ML models are most commonly trained with label supervision, we need a way to extract labels on edge to make FL viable. In this work, we propose a strategy for training FL models using positive and negative user feedback. We also design a novel framework to study different noise patterns in user feedback, and explore how well standard noise-robust objectives can help mitigate this noise when training models in a federated setting. We evaluate our proposed training setup through detailed experiments on two text classification datasets and analyze the effects of varying levels of user reliability and feedback noise on model performance. We show that our method improves substantially over a self-training baseline, achieving performance closer to models trained with full supervision.

cs.LG↗

Training Mixed-Domain Translation Models via Federated Learning

Training mixed-domain translation models is a complex task that demands tailored architectures and costly data preparation techniques. In this work, we leverage federated learning (FL) in order to tackle the problem. Our investigation demonstrates that with slight modifications in the training process, neural machine translation (NMT) engines can be easily adapted when an FL-based aggregation is applied to fuse different domains. Experimental results also show that engines built via FL are able to perform on par with state-of-the-art baselines that rely on centralized training techniques. We evaluate our hypothesis in the presence of five datasets with different sizes, from different domains, to translate from German into English and discuss how FL and NMT can mutually benefit from each other. In addition to providing benchmarking results on the union of FL and NMT, we also propose a novel technique to dynamically control the communication bandwidth by selecting impactful parameters during FL updates. This is a significant achievement considering the large size of NMT engines that need to be exchanged between FL parties.

cs.CL↗