Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 91 records · Page 5Linked to original sources

Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models

Data is a crucial element in large language model (LLM) alignment. Recent studies have explored using LLMs for efficient data collection. However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints. To address these problems, we propose Data Advisor, an enhanced LLM-based method for generating data that takes into account the characteristics of the desired dataset. Starting from a set of pre-defined principles in hand, Data Advisor monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly. Data Advisor can be easily integrated into existing data generation methods to enhance data quality and coverage. Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of Data Advisor in enhancing model safety against various fine-grained safety issues without sacrificing model utility.

cs.CL↗

Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification

We propose a constraint learning schema for fine-tuning Large Language Models (LLMs) with attribute control. Given a training corpus and control criteria formulated as a sequence-level constraint on model outputs, our method fine-tunes the LLM on the training corpus while enhancing constraint satisfaction with minimal impact on its utility and generation quality. Specifically, our approach regularizes the LLM training by penalizing the KL divergence between the desired output distribution, which satisfies the constraints, and the LLM's posterior. This regularization term can be approximated by an auxiliary model trained to decompose the sequence-level constraints into token-level guidance, allowing the term to be measured by a closed-form formulation. To further improve efficiency, we design a parallel scheme for concurrently updating both the LLM and the auxiliary model. We evaluate the empirical performance of our approach by controlling the toxicity when training an LLM. We show that our approach leads to an LLM that produces fewer inappropriate responses while achieving competitive performance on benchmarks and a toxicity detection task.

cs.CL↗

The core collapse of a 16.5 M$_{\odot}$ star

We investigate the 1D stellar evolution of a 16.5 M$_{\odot}$ zero-age main-sequence star having different initial rotations. Starting from the pre-main-sequence, the models evolve up to the onset of the core collapse stage. The collapse of such a massive star can result in several kinds of energetic transients, such as Gamma-Ray Bursts (GRBs), Supernovae, etc. Using the simulation parameters, we calculate their free-fall timescales when the models reach the stage of the onset of core collapse. Estimating the free-fall timescale is crucial for understanding the duration for which the central engine can be fueled, allowing us to compare the free-fall timescale with the T$_{\rm 90}$ duration of GRBs. Our results indicate that, given the constraints of the parameters and initial conditions in our models, rapidly rotating massive stars might serve as potential progenitors of Ultra-Long GRBs (T$_{\rm 90}$ $>>$ 500 sec). In contrast, the non-rotating or slowly rotating models are more prone to explode as hydrogen-rich Type IIP-like core-collapse supernovae.

astro-ph.HE↗

A detailed time-resolved and energy-resolved spectro-polarimetric study of bright GRBs detected by AstroSat CZTI in its first year of operation

The radiation mechanism underlying the prompt emission remains unresolved and can be resolved using a systematic and uniform time-resolved spectro-polarimetric study. In this paper, we investigated the spectral, temporal, and polarimetric characteristics of five bright GRBs using archival data from AstroSat CZTI, Swift BAT, and Fermi GBM. These bright GRBs were detected by CZTI in its first year of operation, and their average polarization characteristics have been published in Chattopadhyay et al. (2022). In the present work, we examined the time-resolved (in 100-600 keV) and energy-resolved polarization measurements of these GRBs with an improved polarimetric technique such as increasing the effective area and bandwidth (by using data from low-gain pixels), using an improved event selection logic to reduce noise in the double events and extend the spectral bandwidth. In addition, we also separately carried out detailed time-resolved spectral analyses of these GRBs using empirical and physical synchrotron models. By these improved time-resolved and energy-resolved spectral and polarimetric studies (not fully coupled spectro-polarimetric fitting), we could pin down the elusive prompt emission mechanism of these GRBs. Our spectro-polarimetric analysis reveals that GRB 160623A, GRB 160703A, and GRB 160821A have Poynting flux-dominated jets. On the other hand, GRB 160325A and GRB 160802A have baryonic-dominated jets with mild magnetization. Furthermore, we observe a rapid change in polarization angle by $\sim$ 90 degrees within the main pulse of very bright GRB 160821A, consistent with our previous results. Our study suggests that the jet composition of GRBs may exhibit a wide range of magnetization, which can be revealed by utilizing spectro-polarimetric investigations of the bright GRBs.

astro-ph.HE↗

Tree-of-Traversals: A Zero-Shot Reasoning Algorithm for Augmenting Black-box Language Models with Knowledge Graphs

Knowledge graphs (KGs) complement Large Language Models (LLMs) by providing reliable, structured, domain-specific, and up-to-date external knowledge. However, KGs and LLMs are often developed separately and must be integrated after training. We introduce Tree-of-Traversals, a novel zero-shot reasoning algorithm that enables augmentation of black-box LLMs with one or more KGs. The algorithm equips a LLM with actions for interfacing a KG and enables the LLM to perform tree search over possible thoughts and actions to find high confidence reasoning paths. We evaluate on two popular benchmark datasets. Our results show that Tree-of-Traversals significantly improves performance on question answering and KG question answering tasks. Code is available at \url{https://github.com/amazon-science/tree-of-traversals}

cs.AI↗

Exploring Origin of Ultra-Long Gamma-ray Bursts: Lessons from GRB 221009A

The brightest Gamma-ray burst (GRB) ever, GRB 221009A, displays ultra-long GRB (ULGRB) characteristics, with a prompt emission duration exceeding 1000 s. To constrain the origin and central engine of this unique burst, we analyze its prompt and afterglow characteristics and compare them to the established set of similar GRBs. To achieve this, we statistically examine a nearly complete sample of Swift-detected GRBs with measured redshifts. Categorizing the sample to Bronze, Silver, and Gold by fitting a Gaussian function to the log-normal of T$_{90}$ duration distribution and considering three sub-samples respectively to 1, 2, and 3 times of the standard deviation to the mean value. GRB 221009A falls into the Gold sub-sample. Our analysis of prompt emission and afterglow characteristics aims to identify trends between the three burst groups. Notably, the Gold sub-sample (a higher likelihood of being ULGRB candidates) suggests a collapsar scenario with a hyper-accreting black hole as a potential central engine, while a few GRBs (GRB 060218, GRB 091024A, and GRB 100316D) in our Gold sub-sample favor a magnetar. Late-time near-IR (NIR) observations from 3.6m Devasthal Optical Telescope (DOT) rule out the presence of any bright supernova associated with GRB 221009A in the Gold sub-sample. To further constrain the physical properties of ULGRB progenitors, we employ the tool MESA to simulate the evolution of low-metallicity massive stars with different initial rotations. The outcomes suggest that rotating ($Ω\geq 0.2\,Ω_{\rm c}$) massive stars could potentially be the progenitors of ULGRBs within the considered parameters and initial inputs to MESA.

astro-ph.HE↗

Magnetar central engine powering the energetic GRB 210610B ?

The bright GRB 210610B was discovered simultaneously by Fermi and Swift missions at redshift 1.13. We utilized broadband Fermi-GBM observations to perform a detailed prompt emission spectral analysis and to understand the radiation physics of the burst. Our analysis displayed that the low energy spectral index ($α_{\rm pt}$) exceeds boundaries expected from the typical synchrotron emission spectrum (-1.5,-0.67), suggesting additional emission signature. We added an additional thermal model with the typical Band or CPL function and found that CPL + BB function is better fitting to the data, suggesting a hybrid jet composition for the burst. Further, we found that the beaming corrected energy (E$_{\rm γ, θ_{j}}$ = 1.06 $\times$ 10$^{51}$ erg) of the burst is less than the total energy budget of the magnetar. Additionally, the X-ray afterglow light curve of this burst exhibits achromatic plateaus, adding another layer of complexity to the explosion's behavior. Interestingly, we noted that the X-ray energy release during the plateau phase (E$_{\rm X,iso}$ = 1.94 $\times$ 10$^{51}$ erg) is also less than the total energy budget of the magnetar. Our results indicate the possibility that a magnetar could be the central engine for this burst.

astro-ph.HE↗

TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality

Automatic Speech Understanding (ASU) aims at human-like speech interpretation, providing nuanced intent, emotion, sentiment, and content understanding from speech and language (text) content conveyed in speech. Typically, training a robust ASU model relies heavily on acquiring large-scale, high-quality speech and associated transcriptions. However, it is often challenging to collect or use speech data for training ASU due to concerns such as privacy. To approach this setting of enabling ASU when speech (audio) modality is missing, we propose TI-ASU, using a pre-trained text-to-speech model to impute the missing speech. We report extensive experiments evaluating TI-ASU on various missing scales, both multi- and single-modality settings, and the use of LLMs. Our findings show that TI-ASU yields substantial benefits to improve ASU in scenarios where even up to 95% of training speech is missing. Moreover, we show that TI-ASU is adaptive to dropout training, improving model robustness in addressing missing speech during inference.

cs.SD↗

Toward Informal Language Processing: Knowledge of Slang in Large Language Models

Recent advancement in large language models (LLMs) has offered a strong potential for natural language systems to process informal language. A representative form of informal language is slang, used commonly in daily conversations and online social media. To date, slang has not been comprehensively evaluated in LLMs due partly to the absence of a carefully designed and publicly accessible benchmark. Using movie subtitles, we construct a dataset that supports evaluation on a diverse set of tasks pertaining to automatic processing of slang. For both evaluation and finetuning, we show the effectiveness of our dataset on two core applications: 1) slang detection, and 2) identification of regional and historical sources of slang from natural sentences. We also show how our dataset can be used to probe the output distributions of LLMs for interpretive insights. We find that while LLMs such as GPT-4 achieve good performance in a zero-shot setting, smaller BERT-like models finetuned on our dataset achieve comparable performance. Furthermore, we show that our dataset enables finetuning of LLMs such as GPT-3.5 that achieve substantially better performance than strong zero-shot baselines. Our work offers a comprehensive evaluation and a high-quality benchmark on English slang based on the OpenSubtitles corpus, serving both as a publicly accessible resource and a platform for applying tools for informal language processing.

cs.CL↗

Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies

Gender-inclusive NLP research has documented the harmful limitations of gender binary-centric large language models (LLM), such as the inability to correctly use gender-diverse English neopronouns (e.g., xe, zir, fae). While data scarcity is a known culprit, the precise mechanisms through which scarcity affects this behavior remain underexplored. We discover LLM misgendering is significantly influenced by Byte-Pair Encoding (BPE) tokenization, the tokenizer powering many popular LLMs. Unlike binary pronouns, BPE overfragments neopronouns, a direct consequence of data scarcity during tokenizer training. This disparate tokenization mirrors tokenizer limitations observed in multilingual and low-resource NLP, unlocking new misgendering mitigation strategies. We propose two techniques: (1) pronoun tokenization parity, a method to enforce consistent tokenization across gendered pronouns, and (2) utilizing pre-existing LLM pronoun knowledge to improve neopronoun proficiency. Our proposed methods outperform finetuning with standard BPE, improving neopronoun accuracy from 14.1% to 58.4%. Our paper is the first to link LLM misgendering to tokenization and deficient neopronoun grammar, indicating that LLMs unable to correctly treat neopronouns as pronouns are more prone to misgender.

cs.CL↗

Harnessing Orbital Hall Effect in Spin-Orbit Torque MRAM

Spin-Orbit Torque (SOT) Magnetic Random-Access Memory (MRAM) devices offer improved power efficiency, nonvolatility, and performance compared to static RAM, making them ideal, for instance, for cache memory applications. Efficient magnetization switching, long data retention, and high-density integration in SOT MRAM require ferromagnets (FM) with perpendicular magnetic anisotropy (PMA) combined with large torques enhanced by Orbital Hall Effect (OHE). We have engineered PMA [Co/Ni]$_3$ FM on selected OHE layers (Ru, Nb, Cr) and investigated the potential of theoretically predicted larger orbital Hall conductivity (OHC) to quantify the torque and switching current in OHE/[Co/Ni]$_3$ stacks. Our results demonstrate a $\sim$30\% enhancement in damping-like torque efficiency with a positive sign for the Ru OHE layer compared to a pure Pt, accompanied by a $\sim$20\% reduction in switching current for Ru compared to pure Pt across more than 250 devices, leading to more than a 60\% reduction in switching power. These findings validate the application of Ru in devices relevant to industrial contexts, supporting theoretical predictions regarding its superior OHC. This investigation highlights the potential of enhanced orbital torques to improve the performance of orbital-assisted SOT-MRAM, paving the way for next-generation memory technology.

physics.app-ph↗

On the steerability of large language models toward data-driven personas

Large language models (LLMs) are known to generate biased responses where the opinions of certain groups and populations are underrepresented. Here, we present a novel approach to achieve controllable generation of specific viewpoints using LLMs, that can be leveraged to produce multiple perspectives and to reflect the diverse opinions. Moving beyond the traditional reliance on demographics like age, gender, or party affiliation, we introduce a data-driven notion of persona grounded in collaborative filtering, which is defined as either a single individual or a cohort of individuals manifesting similar views across specific inquiries. As individuals in the same demographic group may have different personas, our data-driven persona definition allows for a more nuanced understanding of different (latent) social groups present in the population. In addition to this, we also explore an efficient method to steer LLMs toward the personas that we define. We show that our data-driven personas significantly enhance model steerability, with improvements of between $57\%-77\%$ over our best performing baselines.

cs.CL↗

Skyrmionic device for three dimensional magnetic field sensing enabled by spin-orbit torques

Magnetic skyrmions are topologically protected local magnetic solitons that are promising for storage, logic or general computing applications. In this work, we demonstrate that we can use a skyrmion device based on [W/CoFeB/MgO] 1 0 multilayers for three-dimensional magnetic field sensing enabled by spin-orbit torques (SOT). We stabilize isolated chiral skyrmions and stripe domains in the multilayers, as shown by magnetic force microscopy images and micromagnetic simulations. We perform magnetic transport measurements to show that we can sense both in-plane and out-of-plane magnetic fields by means of a differential measurement scheme in which the symmetry of the SOT leads to cancelation of the DC offset. With the magnetic parameters obtained by vibrating sample magnetometry and ferromagnetic resonance measurements, we perform finite-temperature micromagnetic simulations, where we investigate the fundamental origin of the sensing signal. We identify the topological transformation between skyrmions, stripes and type-II bubbles that leads to a change in the resistance that is read-out by the anomalous Hall effect. Our study presents a novel application for skyrmions, where a differential measurement sensing concept is applied to quantify external magnetic fields paving the way towards more energy efficient applications in skyrmionics based spintronics.

cond-mat.mes-hall↗

Partial Federated Learning

Federated Learning (FL) is a popular algorithm to train machine learning models on user data constrained to edge devices (for example, mobile phones) due to privacy concerns. Typically, FL is trained with the assumption that no part of the user data can be egressed from the edge. However, in many production settings, specific data-modalities/meta-data are limited to be on device while others are not. For example, in commercial SLU systems, it is typically desired to prevent transmission of biometric signals (such as audio recordings of the input prompt) to the cloud, but egress of locally (i.e. on the edge device) transcribed text to the cloud may be possible. In this work, we propose a new algorithm called Partial Federated Learning (PartialFL), where a machine learning model is trained using data where a subset of data modalities or their intermediate representations can be made available to the server. We further restrict our model training by preventing the egress of data labels to the cloud for better privacy, and instead use a contrastive learning based model objective. We evaluate our approach on two different multi-modal datasets and show promising results with our proposed approach.

cs.LG↗

Multiwavelength Observations of Gamma Ray Bursts

Gamma-ray bursts (GRBs) are fascinating sources studied in modern astronomy. They are extremely luminous electromagnetic explosions in the Universe observed from cosmological distances. These unique characteristics provide a marvellous chance to study the evolution of massive stars and probe the rarely explored early Universe. In addition, the central source's compactness and the high bulk Lorentz factor in GRB's ultra-relativistic jets make them efficient laboratories for studying high-energy astrophysics. GRBs are the only astrophysical sources observed in two distinct signals: gravitational and electromagnetic waves. GRBs are believed to be produced from a "fireball" moving at a relativistic speed, launched by a fast-rotating black hole or magnetar. GRBs emit radiation in two phases: the initial gamma/hard X-rays prompt emission, the duration of which ranges from a few seconds to hours, followed by the multi-wavelength and long-lived afterglow phase. Based on the observed time frame of GRB prompt emission, astronomers have generally categorized GRBs into two groups: long (> 2 s) and short (< 2 s) bursts. Despite the discovery of GRBs in the late 1960s, their origin is still a great mystery. There are several open questions related to GRBs, such as: What powers the GRBs jets/central engine? What are the possible progenitors? What is the jet composition? What is the underlying emission process that gives rise to observed radiation? Where and how does the energy dissipation occur in the outflow? How to solve the radiative efficiency problem? What are the possible causes of Dark GRBs and orphan afterglows? How to investigate the local environment of GRBs? etc. In this thesis, we explored some of these open enigmas (progenitor, emission mechanisms, jet composition and environment) using multi-wavelength observations obtained using space and ground-based facilities.

astro-ph.HE↗

Faithful Model Evaluation for Model-Based Metrics

Statistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they reflect a genuine relationship. A key step in significance testing is the estimation of confidence interval which is a function of sample variance. Sample variance calculation is straightforward when evaluating against ground truth. However, in many cases, a metric model is often used for evaluation. For example, to compare toxicity of two large language models, a toxicity classifier is used for evaluation. Existing works usually do not consider the variance change due to metric model errors, which can lead to wrong conclusions. In this work, we establish the mathematical foundation of significance testing for model-based metrics. With experiments on public benchmark datasets and a production system, we show that considering metric model errors to calculate sample variances for model-based metrics changes the conclusions in certain experiments.

cs.CL↗

Quantifying the Uncertainty of Sensitivity Coefficients Computed from Uncertain Compound Admittance Matrix and Noisy Grid Measurements

The power-flow sensitivity coefficients (PFSCs) are widely used in the power system for expressing linearized dependencies between the controlled (i.e., the nodal voltages, lines currents) and control variables (e.g., active and reactive power injections, transformer tap positions, etc.). The PFSCs are often computed by knowing the compound admittance matrix of a given network and the grid states. However, when the branch parameters (or admittance matrix) are inaccurate or known with limited accuracy, the computed PFSCs can not be relied upon. Uncertain PFSCs, when used in control, can lead to infeasible control set-points. In this context, this paper presents a method to quantify the uncertainty of the PFSCs from uncertain branch parameters and noisy grid-state measurements that can be used for formulating safe control schemes. We derive an analytical expression using the error-propagation principle. The developed tool is numerically validated using Monte Carlo simulations.

eess.SY↗

Measurement-based/Model-less Estimation of Voltage Sensitivity Coefficients by Feedforward and LSTM Neural Networks in Power Distribution Grids

The increasing adoption of measurement units in electrical power distribution grids has enabled the deployment of data-driven and measurement-based control schemes. Such schemes rely on measurement-based estimated models, where the models are first estimated using raw measurements and then used in the control problem. This work focuses on measurement-based estimation of the voltage sensitivity coefficients which can be used for voltage control. In the existing literature, these coefficients are estimated using regression-based methods, which do not perform well in the case of high measurement noise. This work proposes tackling this problem by using neural network (NN)-based estimation of the voltage sensitivity coefficients which is robust against measurement noise. In particular, we propose using Feedforward and Long-Short Term Memory (LSTM) neural networks. The trained NNs take measurements of nodal voltage magnitudes and active and reactive powers and output the vector of voltage magnitude sensitivity coefficients. The performance of the proposed scheme is compared against the regression-based method for a CIGRE benchmark network.

eess.SY↗