Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 145 records · Page 8Linked to original sources

The long-active afterglow of GRB 210204A: Detection of the most delayed flares in a Gamma-Ray Burst

We present results from extensive broadband follow-up of GRB 210204A over the period of thirty days. We detect optical flares in the afterglow at 7.6 x 10^5 s and 1.1 x 10^6 s after the burst: the most delayed flaring ever detected in a GRB afterglow. At the source redshift of 0.876, the rest-frame delay is 5.8 x 10^5 s (6.71 d). We investigate possible causes for this flaring and conclude that the most likely cause is a refreshed shock in the jet. The prompt emission of the GRB is within the range of typical long bursts: it shows three disjoint emission episodes, which all follow the typical GRB correlations. This suggests that GRB 210204A might not have any special properties that caused late-time flaring, and the lack of such detections for other afterglows might be resulting from the paucity of late-time observations. Systematic late-time follow-up of a larger sample of GRBs can shed more light on such afterglow behaviour. Further analysis of the GRB 210204A shows that the late time bump in the light curve is highly unlikely due to underlying SNe at redshift (z) = 0.876 and is more likely due to the late time flaring activity. The cause of this variability is not clearly quantifiable due to the lack of multi-band data at late time constraints by the bad weather conditions. The flare of GRB 210204A is the latest flare detected to date.

astro-ph.HE↗

Coordinated Day-ahead Dispatch of Multiple Power Distribution Grids hosting Stochastic Resources: An ADMM-based Framework

This work presents an optimization framework to aggregate the power and energy flexibilities in an interconnected power distribution systems. The aggregation framework is used to compute the day-ahead dispatch plans of multiple and interconnected distribution grids operating at different voltage levels. Specifically, the proposed framework optimizes the dispatch plan of an upstream medium voltage (MV) grid accounting for the flexibility offered by downstream low voltage (LV) grids and the knowledge of the uncertainties of the stochastic resources. The framework considers grid, i.e., operational limits on the nodal voltages, lines, and transformer capacity using a linearized grid model, and controllable resources' constraints. The dispatching problem is formulated as a stochastic-optimization scheme considering uncertainty on stochastic power generation and demands and the voltage imposed by the upstream grid. The problem is solved by a distributed optimization method relying on the Alternating Direction Method of Multipliers (ADMM) that splits the main problem into an aggregator problem (solved at the MV-grid level) and several local problems (solved at the MV-connected-controllable-resources and LV-grid levels). The use of distributed optimization enables a decentralized dispatch computation where the centralized aggregator is agnostic about the parameters/models of the participating resources and downstream grids. The framework is validated for interconnected CIGRE medium- and low-voltage networks hosting heterogeneous stochastic and controllable resources.

eess.SY↗

Idele class groups with modulus

We prove Bloch's formula for the Chow group of 0-cycles with modulus on smooth projective varieties over finite fields. The proof relies on two new results in global ramification theory.

math.AG↗

Canary Extraction in Natural Language Understanding Models

Natural Language Understanding (NLU) models can be trained on sensitive information such as phone numbers, zip-codes etc. Recent literature has focused on Model Inversion Attacks (ModIvA) that can extract training data from model parameters. In this work, we present a version of such an attack by extracting canaries inserted in NLU training data. In the attack, an adversary with open-box access to the model reconstructs the canaries contained in the model's training set. We evaluate our approach by performing text completion on canaries and demonstrate that by using the prefix (non-sensitive) tokens of the canary, we can generate the full canary. As an example, our attack is able to reconstruct a four digit code in the training dataset of the NLU model with a probability of 0.5 in its best configuration. As countermeasures, we identify several defense mechanisms that, when combined, effectively eliminate the risk of ModIvA in our experiments.

cs.CL↗

On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations

Multiple metrics have been introduced to measure fairness in various natural language processing tasks. These metrics can be roughly categorized into two categories: 1) \emph{extrinsic metrics} for evaluating fairness in downstream applications and 2) \emph{intrinsic metrics} for estimating fairness in upstream contextualized language representation models. In this paper, we conduct an extensive correlation study between intrinsic and extrinsic metrics across bias notions using 19 contextualized language models. We find that intrinsic and extrinsic metrics do not necessarily correlate in their original setting, even when correcting for metric misalignments, noise in evaluation datasets, and confounding factors such as experiment configuration for extrinsic metrics. %al

cs.CL↗

Mitigating Gender Bias in Distilled Language Models via Counterfactual Role Reversal

Language models excel at generating coherent text, and model compression techniques such as knowledge distillation have enabled their use in resource-constrained settings. However, these models can be biased in multiple ways, including the unfounded association of male and female genders with gender-neutral professions. Therefore, knowledge distillation without any fairness constraints may preserve or exaggerate the teacher model's biases onto the distilled model. To this end, we present a novel approach to mitigate gender disparity in text generation by learning a fair model during knowledge distillation. We propose two modifications to the base knowledge distillation based on counterfactual role reversal$\unicode{x2014}$modifying teacher probabilities and augmenting the training set. We evaluate gender polarity across professions in open-ended text generated from the resulting distilled and finetuned GPT$\unicode{x2012}$2 models and demonstrate a substantial reduction in gender disparity with only a minor compromise in utility. Finally, we observe that language models that reduce gender polarity in language generation do not improve embedding fairness or downstream classification fairness.

cs.CL↗

Measuring Fairness of Text Classifiers via Prediction Sensitivity

With the rapid growth in language processing applications, fairness has emerged as an important consideration in data-driven solutions. Although various fairness definitions have been explored in the recent literature, there is lack of consensus on which metrics most accurately reflect the fairness of a system. In this work, we propose a new formulation : ACCUMULATED PREDICTION SENSITIVITY, which measures fairness in machine learning models based on the model's prediction sensitivity to perturbations in input features. The metric attempts to quantify the extent to which a single prediction depends on a protected attribute, where the protected attribute encodes the membership status of an individual in a protected group. We show that the metric can be theoretically linked with a specific notion of group fairness (statistical parity) and individual fairness. It also correlates well with humans' perception of fairness. We conduct experiments on two text classification datasets : JIGSAW TOXICITY, and BIAS IN BIOS, and evaluate the correlations between metrics and manual annotations on whether the model produced a fair outcome. We observe that the proposed fairness metric based on prediction sensitivity is statistically significantly more correlated with human annotation than the existing counterfactual fairness metric.

cs.LG↗

An Efficient DP-SGD Mechanism for Large Scale NLP Models

Recent advances in deep learning have drastically improved performance on many Natural Language Understanding (NLU) tasks. However, the data used to train NLU models may contain private information such as addresses or phone numbers, particularly when drawn from human subjects. It is desirable that underlying models do not expose private information contained in the training data. Differentially Private Stochastic Gradient Descent (DP-SGD) has been proposed as a mechanism to build privacy-preserving models. However, DP-SGD can be prohibitively slow to train. In this work, we propose a more efficient DP-SGD for training using a GPU infrastructure and apply it to fine-tuning models based on LSTM and transformer architectures. We report faster training times, alongside accuracy, theoretical privacy guarantees and success of Membership inference attacks for our models and observe that fine-tuning with proposed variant of DP-SGD can yield competitive models without significant degradation in training time and improvement in privacy protection. We also make observations such as looser theoretical $ε, δ$ can translate into significant practical privacy gains.

cs.CL↗

Learnings from Federated Learning in the Real world

Federated Learning (FL) applied to real world data may suffer from several idiosyncrasies. One such idiosyncrasy is the data distribution across devices. Data across devices could be distributed such that there are some "heavy devices" with large amounts of data while there are many "light users" with only a handful of data points. There also exists heterogeneity of data across devices. In this study, we evaluate the impact of such idiosyncrasies on Natural Language Understanding (NLU) models trained using FL. We conduct experiments on data obtained from a large scale NLU system serving thousands of devices and show that simple non-uniform device selection based on the number of interactions at each round of FL training boosts the performance of the model. This benefit is further amplified in continual FL on consecutive time periods, where non-uniform sampling manages to swiftly catch up with FL methods using all data at once.

cs.LG↗

GRB 210217A: A short or a long GRB?

Gamma-ray bursts are traditionally classified as short and long bursts based on their $T_{\rm 90}$ value (the time interval during which an instrument observes $5\%$ to $95\%$ of gamma-ray/hard X-ray fluence). However, $T_{\rm 90}$ is dependent on the detector sensitivity and the energy range in which the instrument operates. As a result, different instruments provide different values of $T_{\rm 90}$ for a burst. GRB 210217A is detected with different duration by {\it Swift} and {\it Fermi}. It is classified as a long/soft GRB by {\it Swift}-BAT with a $T_{\rm 90}$ value of 3.76 sec. On the other hand, the sub-threshold detection by {\it Fermi}-GBM classified GRB 210217A as a short/hard burst with a duration of 1.024 sec. We present the multi-wavelength analysis of GRB 210217A (lying in the overlapping regime of long and short GRBs) to identify its actual class using multi-wavelength data. We utilized the $T_{\rm 90}$-hardness ratio, $T_{\rm 90}$-\Ep, and $T_{\rm 90}$-$t_{\rm mvts}$ distributions of the GRBs to find the probability of GRB 210217A being a short GRB. Further, we estimated the photometric redshift of the burst by fitting the joint XRT/UVOT SED and place the burst in the Amati plane. We found that GRB 210217A is an ambiguous burst showing properties of both short and long class of GRBs.

astro-ph.HE↗

Model-less Robust Voltage Control in Active Distribution Networks using Sensitivity Coefficients Estimated from Measurements

Measurement-rich power distribution networks may enable distribution system operators (DSOs) to adopt model-less and measurement-based monitoring and control of distributed energy resources (DERs) for mitigating grid issues such as over/under voltages and lines congestions. However, measurement-based monitoring and control applications may lead to inaccurate control decisions due to measurement errors. In particular, estimation models relying on regression-based schemes result in significant errors in the estimates (e.g., nodal voltages) especially for measurement devices with high Instrument Transformer (IT) classes. The consequences are detrimental to control performance since this may lead to infeasible decisions. This work proposes a model-less robust voltage control accounting for the uncertainties of measurement-based estimated voltage sensitivity coefficients. The coefficients and their uncertainties are obtained using a recursive least squares (RLS)-based online estimation, updated whenever new measurements are available. This formulation is applied to control distributed controllable photovoltaic (PV) generation in a distribution network to restrict the voltage within prescribed limits. The proposed scheme is validated by simulating a CIGRE low-voltage system interfacing multiple controllable PV plants.

eess.SY↗

Probing into emission mechanisms of GRB 190530A using time-resolved spectra and polarization studies: Synchrotron Origin?

Multi-pulsed GRB 190530A, detected by the GBM and LAT onboard \fermi, is the sixth most fluent GBM burst detected so far. This paper presents the timing, spectral, and polarimetric analysis of the prompt emission observed using \AstroSat and \fermi to provide insight into the prompt emission radiation mechanisms. The time-integrated spectrum shows conclusive proof of two breaks due to peak energy and a second lower energy break. Time-integrated (55.43 $\pm$ 21.30 \%) as well as time-resolved polarization measurements, made by the Cadmium Zinc Telluride Imager (CZTI) onboard \AstroSat, show a hint of high degree of polarization. The presence of a hint of high degree of polarization and the values of low energy spectral index ($α_{\rm pt}$) do not run over the synchrotron limit for the first two pulses, supporting the synchrotron origin in an ordered magnetic field. However, during the third pulse, $α_{\rm pt}$ exceeds the synchrotron line of death in few bins, and a thermal signature along with the synchrotron component in the time-resolved spectra is observed. Furthermore, we also report the earliest optical observations constraining afterglow polarization using the MASTER (P $<$ 1.3 \%) and the redshift measurement ($z$= 0.9386) obtained with the 10.4m GTC telescopes. The broadband afterglow can be described with a forward shock model for an ISM-like medium with a wide jet opening angle. We determine a circumburst density of $n_{0} \sim$ 7.41, kinetic energy $E_{\rm K} \sim$ 7.24 $\times 10^{54}$ erg, and radiated $γ$-ray energy $E_{\rm γ, iso} \sim$ 6.05 $\times 10^{54}$ erg, respectively.

astro-ph.HE↗

Revealing nature of GRB 210205A, ZTF21aaeyldq (AT2021any), and follow-up observations with the 4K$\times$4K CCD Imager+3.6m DOT

Optical follow-up observations of optical afterglows of gamma-ray bursts are crucial to probe the geometry of outflows, emission mechanisms, energetics, and burst environments. We performed the follow-up observations of GRB 210205A and ZTF21aaeyldq (AT2021any) using the 3.6m Devasthal Optical Telescope (DOT) around one day after the burst to deeper limits due to the longitudinal advantage of the place. This paper presents our analysis of the two objects using data from other collaborative facilities, i.e., 2.2m Calar Alto Astronomical Observatory (CAHA) and other archival data. Our analysis suggests that GRB 210205A is a potential dark burst once compared with the X-ray afterglow data. Also, comparing results with other known and well-studied dark GRBs samples indicate that the reason for the optical darkness of GRB 210205A could either be intrinsic faintness or a high redshift event. Based on our analysis, we also found that ZTF21aaeyldq is the third known orphan afterglow with a measured redshift except for ZTF20aajnksq (AT2020blt) and ZTF19abvizsw (AT2019pim). The multiwavelength afterglow modelling of ZTF21aaeyldq using the afterglowpy package demands a forward shock model for an ISM-like ambient medium with a rather wider jet opening angle. We determine circumburst density of $n_{0}$ = 0.87 cm$^{-3}$, kinetic energy $E_{k}$ = 3.80 $\times 10^{52}$ erg and the afterglow modelling also indicates that ZTF21aaeyldq is observed on-axis ($θ_{obs} < θ_{core}$) and a gamma-ray counterpart was missed by GRBs satellites. Our results emphasize that the 3.6m DOT has a unique capability for deep follow-up observations of similar and other new transients for deeper observations as a part of time-domain astronomy in the future.

astro-ph.HE↗

Towards Realistic Single-Task Continuous Learning Research for NER

There is an increasing interest in continuous learning (CL), as data privacy is becoming a priority for real-world machine learning applications. Meanwhile, there is still a lack of academic NLP benchmarks that are applicable for realistic CL settings, which is a major challenge for the advancement of the field. In this paper we discuss some of the unrealistic data characteristics of public datasets, study the challenges of realistic single-task continuous learning as well as the effectiveness of data rehearsal as a way to mitigate accuracy loss. We construct a CL NER dataset from an existing publicly available dataset and release it along with the code to the research community.

cs.CL↗

Optimal Grid-Forming Control of Battery Energy Storage Systems Providing Multiple Services: Modelling and Experimental Validation

This paper proposes and experimentally validates a joint control and scheduling framework for a grid-forming converter-interfaced BESS providing multiple services to the electrical grid. The framework is designed to dispatch the operation of a distribution feeder hosting heterogeneous prosumers according to a dispatch plan and provide frequency containment reserve and voltage control as additional services. The framework consists of three phases. In the day-ahead scheduling phase, a robust optimization problem is solved to compute the optimal dispatch plan and frequency droop coefficient, accounting for the uncertainty of the aggregated prosumption. In the intra-day phase, a model predictive control algorithm is used to compute the power set-point for the BESS to achieve the tracking of the dispatch plan. Finally, in a real-time stage, the power setpoint originated by the dispatch tracking is converted into a feasible frequency set-point for the grid forming converter by means of a convex optimisation problem accounting for the capability curve of the power converter. The proposed framework is experimentally validated by using a grid-scale 720 kVA/560 kWh BESS connected to a 20 kV distribution feeder of the EPFL hosting stochastic prosumption and PV generation.

eess.SY↗

Strain Engineering of Epitaxial Pt/Fe Spintronic Terahertz Emitter

Spin-based terahertz (THz) emitters, utilizing the inverse spin Hall effect, are ultra-modern sources for the generation of THz electromagnetic radiation. To make a powerful emitter having large THz amplitude and bandwidth, fundamental understanding in terms of microscopic models is essential. This study reveals important factors to engineer the THz emission amplitude and bandwidth in epitaxial Pt/Fe emitters grown on MgO and MgAl$_2$O$_4$ (MAO) substrates, where the choice of the substrate plays an important role. The THz amplitude and bandwidth are affected by the induced strain in the Fe spin source layer. On the one hand, the THz electric field amplitude is found to be larger when Pt/Fe is grown on MgO even though the crystalline quality of the Fe film is superior when grown on MAO. This is because of the larger defect density, smaller electron relaxation time, and lower electrical conductivity in the THz regime when Fe is grown on MgO. On the other hand, the bandwidth is found to be larger for Pt/Fe grown on MAO and is explained by the uncoupled/coupled Lorentz oscillator models. This study provides an insightful pathway to further engineer metallic spintronic THz emitters in terms of the proper choice of substrate and microscopic properties of the emitter layers.

cond-mat.mtrl-sci↗

Motivic invariants of symmetric powers of curves

We study the structure of various invariants of the symmetric powers of a smooth projective curve in terms of that of the Jacobian of the curve. We generalise the results of Macdonald and Collino to various invariants including the Weil-cohomology theory, the higher Chow groups, the additive higher Chow groups and the rational $K$-groups.

math.AG↗

Does Robustness Improve Fairness? Approaching Fairness with Word Substitution Robustness Methods for Text Classification

Existing bias mitigation methods to reduce disparities in model outcomes across cohorts have focused on data augmentation, debiasing model embeddings, or adding fairness-based optimization objectives during training. Separately, certified word substitution robustness methods have been developed to decrease the impact of spurious features and synonym substitutions on model predictions. While their end goals are different, they both aim to encourage models to make the same prediction for certain changes in the input. In this paper, we investigate the utility of certified word substitution robustness methods to improve equality of odds and equality of opportunity on multiple text classification tasks. We observe that certified robustness methods improve fairness, and using both robustness and bias mitigation methods in training results in an improvement in both fronts

cs.CL↗