Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 109 records · Page 6Linked to original sources

JAB: Joint Adversarial Prompting and Belief Augmentation

With the recent surge of language models in different applications, attention to safety and robustness of these models has gained significant importance. Here we introduce a joint framework in which we simultaneously probe and improve the robustness of a black-box target model via adversarial prompting and belief augmentation using iterative feedback loops. This framework utilizes an automated red teaming approach to probe the target model, along with a belief augmenter to generate instructions for the target model to improve its robustness to those adversarial probes. Importantly, the adversarial model and the belief generator leverage the feedback from past interactions to improve the effectiveness of the adversarial prompts and beliefs, respectively. In our experiments, we demonstrate that such a framework can reduce toxic content generation both in dynamic cases where an adversary directly interacts with a target model and static cases where we use a static benchmark dataset to evaluate our model.

cs.AI↗

Strong in-plane magnetic anisotropy (Co0.15Fe0.85)5GeTe2/graphene van der Waals heterostructure spin-valve at room temperature

Van der Waals (vdW) magnets are promising owing to their tunable magnetic properties with doping or alloy composition, where the strength of magnetic interactions, their symmetry, and magnetic anisotropy can be tuned according to the desired application. However, most of the vdW magnet based spintronic devices are so far limited to cryogenic temperatures with magnetic anisotropies favouring out-of-plane or canted orientation of the magnetization. Here, we report room-temperature lateral spin-valve devices with strong in-plane magnetic anisotropy of the vdW ferromagnet (Co0.15Fe0.85)5GeTe2 (CFGT) in heterostructures with graphene. Magnetization measurements reveal above room-temperature ferromagnetism in CFGT with a strong in-plane magnetic anisotropy. Density functional theory calculations show that the magnitude of the anisotropy depends on the Co concentration and is caused by the substitution of Co in the outermost Fe layer. Heterostructures consisting of CFGT nanolayers and graphene were used to experimentally realize basic building blocks for spin valve devices such as efficient spin injection and detection. The spin transport and Hanle spin precession measurements prove a strong in-plane and negative spin polarization at the interface with graphene, which is supported by the calculated spin-polarized density of states of CFGT. The in-plane magnetization of CFGT at room temperature proves its usefulness in graphene lateral spin-valve devices, thus opening further opportunities for spintronic technologies.

cond-mat.mes-hall↗

Evaluating Large Language Models on Controlled Generation Tasks

While recent studies have looked into the abilities of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc, there have been few studies looking into the controllability of large language models on generation tasks. We present an extensive analysis of various benchmarks including a sentence planning benchmark with different granularities. After comparing large language models against state-of-the-start finetuned smaller models, we present a spectrum showing large language models falling behind, are comparable, or exceed the ability of smaller models. We conclude that **large language models struggle at meeting fine-grained hard constraints**.

cs.CL↗

Coordinated Replay Sample Selection for Continual Federated Learning

Continual Federated Learning (CFL) combines Federated Learning (FL), the decentralized learning of a central model on a number of client devices that may not communicate their data, and Continual Learning (CL), the learning of a model from a continual stream of data without keeping the entire history. In CL, the main challenge is \textit{forgetting} what was learned from past data. While replay-based algorithms that keep a small pool of past training data are effective to reduce forgetting, only simple replay sample selection strategies have been applied to CFL in prior work, and no previous work has explored coordination among clients for better sample selection. To bridge this gap, we adapt a replay sample selection objective based on loss gradient diversity to CFL and propose a new relaxation-based selection of samples to optimize the objective. Next, we propose a practical algorithm to coordinate gradient-based replay sample selection across clients without communicating private data. We benchmark our coordinated and uncoordinated replay sample selection algorithms against random sampling-based baselines with language models trained on a large scale de-identified real-world text dataset. We show that gradient-based sample selection methods both boost performance and reduce forgetting compared to random sampling methods, with our coordination method showing gains early in the low replay size regime (when the budget for storing past data is small).

cs.LG↗

Revealing characteristics of dark GRB 150309A: dust extinguished or high-z?

Dark GRBs constitute a significant fraction of the GRB population. In this paper, we present the multiwavelength analysis of an intense two-episodic GRB 150309A observed early on to ~114 days post-burst. Despite the strong gamma-ray emission, no optical afterglow was detected for this burst. However, we discovered near-infrared afterglow ($K_{\rm S}$-band), ~5.2 hours post burst, with the CIRCE instrument mounted at the 10.4m GTC. We used Fermi observations of GRB 150309A to understand the prompt emission mechanisms and jet composition. We performed the early optical observations using the BOOTES robotic telescope and late-time afterglow observations using the GTC. A potential faint host galaxy is also detected at optical wavelength using the GTC. We modelled the potential host galaxy of GRB 150309A in order to explore the environment of the burst. The time-resolved spectral analysis of Fermi data indicates a hybrid jet composition consisting of a matter-dominated fireball and magnetic-dominated Poynting flux. GTC observations of the afterglow revealed that the counterpart of GRB 150309A was very red, with H-$K_{\rm S}$ > 2.1 mag (95 $\%$ confidence). The red counterpart was not discovered in any bluer filters of Swift UVOT, indicative of high redshift origin. This possibility was discarded based on multiple arguments, such as spectral analysis of X-ray afterglow constrain z < 4.15 and a moderate redshift value obtained using spectral energy distribution modelling of the potential galaxy. The broadband afterglow SED implies a very dusty host galaxy with deeply embedded GRB (suggesting $A_{\rm V}$ $\gtrsim$ 35 mag). The environment of GRB 150309A demands a high extinction towards the line of sight, demanding dust obscuration is the most probable origin of optical darkness and the very red afterglow of GRB 150309A. This result makes GRB 150309A the highest extinguished GRB known to date.

astro-ph.HE↗

Quasi-integrability and nonlinear resonances in cold atoms under modulation

Quantum dynamics of a collection of atoms subjected to phase modulation has been carefully revisited. We present an exact analysis of the evolution of a two-level system (represented by a spinor) under the action of a time-dependent matrix Hamiltonian. The dynamics is shown to evolve on two coupled potential energy surfaces, one of them binding while the other one scattering type. The dynamics is shown to be quasi-integrable with nonlinear resonances. The bounded dynamics with intermittent scattering at random moments presents the scenario reminiscent to Anderson and dynamical localization. We believe that a careful analytical investigation of a multi-component system which is classically non-integrable is relevant to many other fields, including quantum computation with multi-qubit system.

quant-ph↗

Recent observations of peculiar Gamma-ray bursts using 3.6 m Devasthal Optical Telescope (DOT)

India has been actively involved in the follow-up observations of optical afterglows of gamma-ray bursts (GRBs) for more than two decades, using the country's meter-class facilities such as the 1.04 m Sampurnanand Telescope, 1.3 m Devasthal Fast Optical Telescope, 2.01 m Himalayan Chandra Telescope along with many others in the country, utilizing the longitudinal advantage of the place. However, since 2016, Indian astronomers have embarked on a new era of exploration by utilizing the country's largest optical telescope, the 3.6 m Devasthal Optical Telescope (DOT) at the Devasthal Observatory of ARIES Nainital. This unique telescope has opened up exciting opportunities for transient study. Starting from the installation itself, the DOT has been actively performing the target of opportunity (ToO) observations, leading to many interesting discoveries. Notable achievements include the contributions towards the discovery of long GRB 211211A arising from a binary merger, the discovery of the most delayed optical flare from GRB 210204A along with the very faint optical afterglow (fainter than 25 mag in g-band) of GRB 200412B. We also successfully observed the optical counterpart of the very-high-energy (VHE) detected burst GRB 201015A using DOT. Additionally, DOT has been used for follow-up observations of dark and orphan afterglows, along with the observations of host galaxies associated with peculiar GRBs. More recently, DOT's near-IR follow-up capabilities helped us to detect the first near-IR counterpart (GRB 230409B) using an Indian telescope. In this work, we summarise the recent discoveries and observations of GRBs using the 3.6 m DOT, highlighting the significant contributions in revealing the mysteries of these cosmic transients.

astro-ph.HE↗

Evolution and Final Fates of a Rotating 25 M$_{\odot}$ Pop III star

In this proceeding, we present the 1-dimensional stellar evolution of two rotating population III (Pop III) star models, each having a mass of 25 M$_{\odot}$ at the zero-age main-sequence (ZAMS). The slowly rotating model has an initial angular rotational velocity of 10 per cent of the critical angular rotational velocity. In contrast, the rapidly rotating model has an initial angular rotational velocity of 70 per cent of the critical angular rotational velocity. As an effect of rotationally enhanced mixing, we find that the rapidly rotating model suffers an enormous mass loss due to the deposition of a significant amount of CNO elements toward the surface after the main-sequence phase. We also display the simulated light curves as these models explode into core-collapse supernovae (CCSNe).

astro-ph.HE↗

4K$\times$4K CCD Imager for the 3.6m DOT: Recent up-gradations and results

The 4K$\times$4K CCD Imager is the first light instrument for the 3.6m Devasthal Optical Telescope and is producing broad-band imaging observations of many Galactic and extra-galactic sources since 2015-2016. Capabilities of the CCD Imager are demonstrated recently through several publications using the well-calibrated multi-band deep photometric results as expected from other similar facilities globally. In this article, we summarize some of the recent up-gradations made to improve the Imager, i.e., mounting the new filter wheel casing, replacing stray light baffles and discussing the fringe pattern corrections in redder filters. Some of the new science initiatives like galaxy-embedded faint point sources including WR stars and the observations of low surface brightness galaxy clusters are also discussed.

astro-ph.IM↗

FedMultimodal: A Benchmark For Multimodal Federated Learning

Over the past few years, Federated Learning (FL) has become an emerging machine learning technique to tackle data privacy challenges through collaborative training. In the Federated Learning algorithm, the clients submit a locally trained model, and the server aggregates these parameters until convergence. Despite significant efforts that have been made to FL in fields like computer vision, audio, and natural language processing, the FL applications utilizing multimodal data streams remain largely unexplored. It is known that multimodal learning has broad real-world applications in emotion recognition, healthcare, multimedia, and social media, while user privacy persists as a critical concern. Specifically, there are no existing FL benchmarks targeting multimodal applications or related tasks. In order to facilitate the research in multimodal FL, we introduce FedMultimodal, the first FL benchmark for multimodal learning covering five representative multimodal applications from ten commonly used datasets with a total of eight unique modalities. FedMultimodal offers a systematic FL pipeline, enabling end-to-end modeling framework ranging from data partition and feature extraction to FL benchmark algorithms and model evaluation. Unlike existing FL benchmarks, FedMultimodal provides a standardized approach to assess the robustness of FL against three common data corruptions in real-life multimodal applications: missing modalities, missing labels, and erroneous labels. We hope that FedMultimodal can accelerate numerous future research directions, including designing multimodal FL algorithms toward extreme data heterogeneity, robustness multimodal FL, and efficient multimodal FL. The datasets and benchmark results can be accessed at: https://github.com/usc-sail/fed-multimodal.

cs.DC↗

"I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation

Transgender and non-binary (TGNB) individuals disproportionately experience discrimination and exclusion from daily life. Given the recent popularity and adoption of language generation technologies, the potential to further marginalize this population only grows. Although a multitude of NLP fairness literature focuses on illuminating and addressing gender biases, assessing gender harms for TGNB identities requires understanding how such identities uniquely interact with societal gender norms and how they differ from gender binary-centric perspectives. Such measurement frameworks inherently require centering TGNB voices to help guide the alignment between gender-inclusive NLP and whom they are intended to serve. Towards this goal, we ground our work in the TGNB community and existing interdisciplinary literature to assess how the social reality surrounding experienced marginalization of TGNB persons contributes to and persists within Open Language Generation (OLG). This social knowledge serves as a guide for evaluating popular large language models (LLMs) on two key aspects: (1) misgendering and (2) harmful responses to gender disclosure. To do this, we introduce TANGO, a dataset of template-based real-world text curated from a TGNB-oriented community. We discover a dominance of binary gender norms reflected by the models; LLMs least misgendered subjects in generated text when triggered by prompts whose subjects used binary pronouns. Meanwhile, misgendering was most prevalent when triggering generation with singular they and neopronouns. When prompted with gender disclosures, TGNB disclosure generated the most stigmatizing language and scored most toxic, on average. Our findings warrant further research on how TGNB harms manifest in LLMs and serve as a broader case study toward concretely grounding the design of gender-inclusive AI in community voices and interdisciplinary literature.

cs.CL↗

Multi-VALUE: A Framework for Cross-Dialectal English NLP

Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect invariant, meaning that performance remains constant over dialectal shifts. Current systems often fall short of this ideal since they are designed and tested on a single dialect: Standard American English (SAE). We introduce a suite of resources for evaluating and achieving English dialect invariance. The resource is called Multi-VALUE, a controllable rule-based translation system spanning 50 English dialects and 189 unique linguistic features. Multi-VALUE maps SAE to synthetic forms of each dialect. First, we use this system to stress tests question answering, machine translation, and semantic parsing. Stress tests reveal significant performance disparities for leading models on non-standard dialects. Second, we use this system as a data augmentation technique to improve the dialect robustness of existing systems. Finally, we partner with native speakers of Chicano and Indian English to release new gold-standard variants of the popular CoQA task. To execute the transformation code, run model checkpoints, and download both synthetic and gold-standard dialectal benchmark datasets, see http://value-nlp.org.

cs.CL↗

Controlling the Extraction of Memorized Data from Large Language Models via Prompt-Tuning

Large Language Models (LLMs) are known to memorize significant portions of their training data. Parts of this memorized content have been shown to be extractable by simply querying the model, which poses a privacy risk. We present a novel approach which uses prompt-tuning to control the extraction rates of memorized content in LLMs. We present two prompt training strategies to increase and decrease extraction rates, which correspond to an attack and a defense, respectively. We demonstrate the effectiveness of our techniques by using models from the GPT-Neo family on a public benchmark. For the 1.3B parameter GPT-Neo model, our attack yields a 9.3 percentage point increase in extraction rate compared to our baseline. Our defense can be tuned to achieve different privacy-utility trade-offs by a user-specified hyperparameter. We achieve an extraction rate reduction of up to 97.7% relative to our baseline, with a perplexity increase of 16.9%.

cs.CL↗

Experimental Validation of Model-less Robust Voltage Control using Measurement-based Estimated Voltage Sensitivity Coefficients

Increasing adoption of smart meters and phasor measurement units (PMUs) in power distribution networks are enabling the adoption of data-driven/model-less control schemes to mitigate grid issues such as over/under voltages and power-flow congestions. However, such a scheme can lead to infeasible/inaccurate control decisions due to measurement inaccuracies. In this context, the authors' previous work proposed a robust measurement-based control scheme accounting for the uncertainties of the estimated models. In this scheme, a recursive least squares (RLS)-based method estimates the grid model (in the form of voltage magnitude sensitivity coefficients). Then, a robust control problem optimizes power set-points of distributed energy resources (DERs) such that the nodal voltage limits are satisfied. The estimated voltage sensitivity coefficients are used to model the nodal voltages, and the control robustness is achieved by accounting for their uncertainties. This work presents the first experimental validation of such a robust model-less control scheme on a real power distribution grid. The scheme is applied for voltage control by regulating two photovoltaic (PV) inverters connected in a real microgrid which is a replica of the CIGRE benchmark microgrid network at the EPFL Distributed Electrical Systems Laboratory.

eess.SY↗

Single device offset-free magnetic field sensing principle with tunable sensitivity and linear range based on spin-orbit-torques

We propose a novel device concept using spin-orbit-torques to realize a magnetic field sensor, where we eliminate the sensor offset using a differential measurement concept. We derive a simple analytical formulation for the sensor signal and demonstrate its validity with numerical investigations using macrospin simulations. The sensitivity and the measurable linear sensing range in the proposed concept can be tuned by either varying the effective magnetic anisotropy or by varying the magnitude of the injected currents. We show that undesired perturbation fields normal to the sensitive direction preserve the zero-offset property and only slightly modulate the sensitivity of the proposed sensor. Higher-harmonics voltage analysis on a Hall cross experimentally confirms the linearity and tunability via current strength. Additionally, the sensor exhibits a non-vanishing offset in the experiment which we attribute to the anomalous Nernst effect.

cond-mat.mes-hall↗

MUTANT: A Multi-sentential Code-mixed Hinglish Dataset

The multi-sentential long sequence textual data unfolds several interesting research directions pertaining to natural language processing and generation. Though we observe several high-quality long-sequence datasets for English and other monolingual languages, there is no significant effort in building such resources for code-mixed languages such as Hinglish (code-mixing of Hindi-English). In this paper, we propose a novel task of identifying multi-sentential code-mixed text (MCT) from multilingual articles. As a use case, we leverage multilingual articles from two different data sources and build a first-of-its-kind multi-sentential code-mixed Hinglish dataset i.e., MUTANT. We propose a token-level language-aware pipeline and extend the existing metrics measuring the degree of code-mixing to a multi-sentential framework and automatically identify MCT in the multilingual articles. The MUTANT dataset comprises 67k articles with 85k identified Hinglish MCTs. To facilitate future research, we make the publicly available.

cs.CL↗

Evolution of Rotating 25 M$_{\odot}$ Population III star: Physical Properties and Resulting Supernovae

In this Letter, we report the outcomes of 1-D modelling of a rotating 25 M$_{\odot}$ zero-age main-sequence Population III star up to the stage of the onset of core collapse. Rapidly rotating models display violent and sporadic mass losses after the Main-Sequence stage. In comparison to the solar metallicity model, Pop III models show very small pre-supernova radii. Further, with models at the stage of the onset of core collapse, we simulate the hydrodynamic simulations of resulting supernovae. Depending upon the mass losses due to corresponding rotations and stellar winds, the resulting supernovae span a class from weak Type II to Type Ib/c. We find that the absolute magnitudes of the core-collapse supernovae resulting from Pop III stars are much fainter than that resulting from a solar metallicity star. From our simulation results, we also conclude that within the considered limits of explosion energies and Nickel masses, these transient events are very faint, making it difficult for them to be detected at high redshifts.

astro-ph.SR↗

Detection of long-range orbital-Hall torques

We report and quantify a large orbital-Hall torque generated by Nb and Ru, which we identify from a strong dependence of torques on the ferromagnets. This is manifested as a sign reversal and strong enhancement in the damping-like torques measured in Nb (or Ru)/Ni bilayers as compared to Nb (or Ru)/FeCoB bilayers. The long-range nature of orbital transport in the ferromagnet is revealed by the thickness dependences of Ni in Nb (or Ru)/Ni bilayers which are markedly different from the regular spin absorption in the ferromagnet that takes place within a few angstroms and thus it uniquely distinguishes the orbital Hall torque from the spin Hall torque.

cond-mat.mes-hall↗