Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 73 records · Page 4Linked to original sources

Kaleidoscopic Teaming in Multi Agent Simulations

Warning: This paper contains content that may be inappropriate or offensive. AI agents have gained significant recent attention due to their autonomous tool usage capabilities and their integration in various real-world applications. This autonomy poses novel challenges for the safety of such systems, both in single- and multi-agent scenarios. We argue that existing red teaming or safety evaluation frameworks fall short in evaluating safety risks in complex behaviors, thought processes and actions taken by agents. Moreover, they fail to consider risks in multi-agent setups where various vulnerabilities can be exposed when agents engage in complex behaviors and interactions with each other. To address this shortcoming, we introduce the term kaleidoscopic teaming which seeks to capture complex and wide range of vulnerabilities that can happen in agents both in single-agent and multi-agent scenarios. We also present a new kaleidoscopic teaming framework that generates a diverse array of scenarios modeling real-world human societies. Our framework evaluates safety of agents in both single-agent and multi-agent setups. In single-agent setup, an agent is given a scenario that it needs to complete using the tools it has access to. In multi-agent setup, multiple agents either compete against or cooperate together to complete a task in the scenario through which we capture existing safety vulnerabilities in agents. We introduce new in-context optimization techniques that can be used in our kaleidoscopic teaming framework to generate better scenarios for safety analysis. Lastly, we present appropriate metrics that can be used along with our framework to measure safety of agents. Utilizing our kaleidoscopic teaming framework, we identify vulnerabilities in various models with respect to their safety in agentic use-cases.

cs.AI↗

Cavity Optomechanical Quantum Memory for Twisted Photons Using a Ring BEC

We theoretically propose a photonic orbital angular momentum (OAM) quantum memory platform based on an atomic Bose-Einstein condensate confined in a ring trap and placed inside a Fabry-Perot cavity driven by Laguerre-Gaussian beams. In contrast to electromagnetically induced transparency-based protocols, our memory does not require change of internal atomic levels. The optical states are instead stored in the large Hilbert space of topologically protected and long-lived motional states (persistent currents) of the condensate, yielding a storage time three orders of magnitude better than presently available. Further, the use of a cavity provides orders of magnitude more resonances, and hence bandwidth, for reading and writing than internal atomic transitions. Finally, the analogy to cavity optomechanics suggests a natural path to wavelength conversion, OAM transduction, and nondestructive readout of the memory.

quant-ph↗

Not Every Token Needs Forgetting: Selective Unlearning to Limit Change in Utility in Large Language Model Unlearning

Large Language Model (LLM) unlearning has recently gained significant attention, driven by the need to remove unwanted information, such as private, sensitive, or copyrighted content, from LLMs. However, conventional unlearning approaches indiscriminately update model parameters to forget all tokens in a target document, including common tokens (e.g., pronouns, prepositions, general nouns) that carry general knowledge. In this paper, we highlight that not every token needs forgetting. We propose Selective Unlearning (SU), which identifies a critical subset of tokens within the forgetting set that is relevant to the unwanted information, and unlearns only those tokens. Experiments on two benchmarks and six baseline unlearning algorithms demonstrate that SU not only achieves effective unlearning on the targeted forget data, but also significantly preserves the model's utility in the retaining set.

cs.CL↗

Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation

Safety reasoning is a recent paradigm where LLMs reason over safety policies before generating responses, thereby mitigating limitations in existing safety measures such as over-refusal and jailbreak vulnerabilities. However, implementing this paradigm is challenging due to the resource-intensive process of creating high-quality policy-embedded chain-of-thought (CoT) datasets while ensuring reasoning remains accurate and free from hallucinations or policy conflicts. To tackle this, we propose AIDSAFE: Agentic Iterative Deliberation for Safety Reasoning, a novel data generation recipe that leverages multi-agent deliberation to iteratively expand reasoning on safety policies. A data refiner stage in AIDSAFE ensures high-quality outputs by eliminating repetitive, redundant, and deceptive thoughts. AIDSAFE-generated CoTs provide a strong foundation for supervised fine-tuning (SFT)-based safety training. Additionally, to address the need of preference data in alignment stages, such as DPO training, we introduce a supplemental recipe that uses belief augmentation to create distinct selected and rejected CoT samples. Our evaluations demonstrate that AIDSAFE-generated CoTs achieve superior policy adherence and reasoning quality. Consequently, we show that fine-tuning open-source LLMs on these CoTs can significantly improve safety generalization and jailbreak robustness while maintaining acceptable utility and over-refusal accuracy. AIDSAFE-generated CoT datasets can be found here: https://huggingface.co/datasets/AmazonScience/AIDSAFE

cs.AI↗

Certifying Counterfactual Bias in LLMs

Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across LLM responses for different demographic groups (a.k.a. counterfactual bias), as they do not scale to large number of inputs and do not provide guarantees. Therefore, we propose the first framework, LLMCert-B that certifies LLMs for counterfactual bias on distributions of prompts. A certificate consists of high-confidence bounds on the probability of unbiased LLM responses for any set of counterfactual prompts - prompts differing by demographic groups, sampled from a distribution. We illustrate counterfactual bias certification for distributions of counterfactual prompts created by applying prefixes sampled from prefix distributions, to a given set of prompts. We consider prefix distributions consisting random token sequences, mixtures of manual jailbreaks, and perturbations of jailbreaks in LLM's embedding space. We generate non-trivial certificates for SOTA LLMs, exposing their vulnerabilities over distributions of prompts generated from computationally inexpensive prefix distributions.

cs.AI↗

Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning

The exploitation of large language models (LLMs) for malicious purposes poses significant security risks as these models become more powerful and widespread. While most existing red-teaming frameworks focus on single-turn attacks, real-world adversaries typically operate in multi-turn scenarios, iteratively probing for vulnerabilities and adapting their prompts based on threat model responses. In this paper, we propose \AlgName, a novel multi-turn red-teaming agent that emulates sophisticated human attackers through complementary learning dimensions: global tactic-wise learning that accumulates knowledge over time and generalizes to new attack goals, and local prompt-wise learning that refines implementations for specific goals when initial attempts fail. Unlike previous multi-turn approaches that rely on fixed strategy sets, \AlgName enables the agent to identify new jailbreak tactics, develop a goal-based tactic selection framework, and refine prompt formulations for selected tactics. Empirical evaluations on JailbreakBench demonstrate our framework's superior performance, achieving over 90\% attack success rates against GPT-3.5-Turbo and Llama-3.1-70B within 5 conversation turns, outperforming state-of-the-art baselines. These results highlight the effectiveness of dynamic learning in identifying and exploiting model vulnerabilities in realistic multi-turn scenarios.

cs.AI↗

SemEval-2025 Task 4: Unlearning sensitive content from Large Language Models

We introduce SemEval-2025 Task 4: unlearning sensitive content from Large Language Models (LLMs). The task features 3 subtasks for LLM unlearning spanning different use cases: (1) unlearn long form synthetic creative documents spanning different genres; (2) unlearn short form synthetic biographies containing personally identifiable information (PII), including fake names, phone number, SSN, email and home addresses, and (3) unlearn real documents sampled from the target model's training dataset. We received over 100 submissions from over 30 institutions and we summarize the key techniques and lessons in this paper.

cs.CL↗

Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base

Large language models (LLMs) possess impressive linguistic capabilities but often fail to faithfully retain factual knowledge, leading to hallucinations and unreliable outputs. Understanding LLMs' knowledge deficiencies by exhaustively evaluating against full-scale knowledge bases is computationally prohibitive, especially for closed-weight models. We propose stochastic error ascent (SEA), a scalable and efficient framework for discovering knowledge deficiencies (errors) in closed-weight LLMs under a strict query budget. Rather than naively probing all knowledge candidates, SEA formulates error discovery as a stochastic optimization process: it iteratively retrieves new high-error candidates by leveraging the semantic similarity to previously observed failures. To further enhance search efficiency and coverage, SEA employs hierarchical retrieval across document and paragraph levels, and constructs a relation directed acyclic graph to model error propagation and identify systematic failure modes. Empirically, SEA uncovers 40.7x more knowledge errors than Automated Capability Discovery and 26.7% more than AutoBencher, while reducing the cost-per-error by 599x and 9x, respectively. Human evaluation confirms the high quality of generated questions, while ablation and convergence analyses validate the contribution of each component in SEA. Further analysis on the discovered errors reveals correlated failure patterns across LLM families and recurring deficits, highlighting the need for better data coverage and targeted fine-tuning in future LLM development.

cs.CL↗

K-Edit: Language Model Editing with Contextual Knowledge Awareness

As the world changes, we need to be able to update our models and correct false information without costly retraining. Knowledge-based model editing enables precise modifications to the weights of large language models in order to modify the information encoded within. Recent approaches have seen success in enabling recall of edited information for thousands of edits at once. However, these approaches fail to produce edits that account for associated contextual information. We present K-Edit, an effective approach to generating contextually consistent knowledge edits. By using knowledge graphs, which maintain contextual consistency when an edge is edited, we are able to generate additional \textit{contextual edits} that ensure consistency of related information in the language model. Our experiments demonstrate significant improvements in multi-hop question answering while maintaining the general effectiveness and scalability of model edits.

cs.LG↗

LUME: LLM Unlearning with Multitask Evaluations

Unlearning aims to remove copyrighted, sensitive, or private content from large language models (LLMs) without a full retraining. In this work, we develop a multi-task unlearning benchmark (LUME) which features three tasks: (1) unlearn synthetically generated creative short novels, (2) unlearn synthetic biographies with sensitive information, and (3) unlearn a collection of public biographies. We further release two fine-tuned LLMs of 1B and 7B parameter sizes as the target models. We conduct detailed evaluations of several recently proposed unlearning algorithms and present results on carefully crafted metrics to understand their behavior and limitations.

cs.CL↗

X-ray and gamma-ray timing of GRB 180720B, GRB 181222B, GRB 211211A and GRB 220910A observed with Fermi and ASIM

We present a timing study of the gamma and X-ray observations and analysis of a sample of bright gamma-ray bursts (GRBs; i.e. GRB 180720B, GRB 181222B, GRB 211211A and GRB 220910A), including the very bright and long GRB 211211A (a.k.a. kilonova candidate). They have been detected and observed by the Atmosphere-Space Interactions Monitor (ASIM) installed on the International Space Station (ISS) and the Gamma-ray Burst Monitor (GBM) on-board the Fermi mission. The early (T-T0=s) and high-energy (0.3-20 MeV) ASIM High Energy Detector (HED) and (150 keV-30 MeV) Fermi (BGO) light curves show well-defined peaks with a low quasi-periodic oscillation (QPO) frequency between 2.5-3.5 Hz that could be identified with the spin of the neutron star in the binary mergers (coinciding with the orbital frequency of the binary merger) originating these GRBs. These QPOs consist on the first detection of low-frequency QPOs (<10 Hz) detected in magnetars so far. We also detect a strong QPO at 21.8-22 Hz in GRB 181222B together with its (less significant) harmonics. The low-frequency QPO would correspond to the signal of the orbiting neutron star (NS) previous to the final coalescence giving rise to the gravitational-wave (GW) signal.

astro-ph.HE↗

Identification of orbital pumping from spin pumping and rectification effects

The recently predicted mechanism of orbital pumping enables the generation of pure orbital current from a precessing ferromagnet (FM) without the need for electrical current injection. This orbital current can be efficiently injected into an adjacent nonmagnetic material (NM) without being hampered by electrical conductivity mismatch. However, experimentally identifying this novel effect presents significant challenges due to the substantial background contributions from spin pumping and spin rectification effects (SREs). In this work, we disentangle the effects of orbital pumping from spin pumping in bilayer structures composed of Nb/Ni and Nb/$\mathrm{Fe_{60}Co_{20}B_{20}}$ by observing a sign reversal of the measured voltage. This reversal arises from the competing signs of the spin and orbital Hall effects in the Nb. We establish methods to differentiate the pumping signal from SREs by analyzing the distinct angular dependence of the measured voltage and its spatial dependence relative to the radio frequency excitation source.

cond-mat.mtrl-sci↗

Embodied Red Teaming for Auditing Robotic Foundation Models

Language-conditioned robot models have the potential to enable robots to perform a wide range of tasks based on natural language instructions. However, assessing their safety and effectiveness remains challenging because it is difficult to test all the different ways a single task can be phrased. Current benchmarks have two key limitations: they rely on a limited set of human-generated instructions, missing many challenging cases, and focus only on task performance without assessing safety, such as avoiding damage. To address these gaps, we introduce Embodied Red Teaming (ERT), a new evaluation method that generates diverse and challenging instructions to test these models. ERT uses automated red teaming techniques with Vision Language Models (VLMs) to create contextually grounded, difficult instructions. Experimental results show that state-of-the-art language-conditioned robot models fail or behave unsafely on ERT-generated instructions, underscoring the shortcomings of current benchmarks in evaluating real-world performance and safety. Code and videos are available at: https://s-karnik.github.io/embodied-red-team-project-page.

cs.RO↗

Sensing atomic superfluid rotation beyond the standard quantum limit

Atomic superfluids formed using Bose-Einstein condensates (BECs) in a ring trap are currently being investigated in the context of superfluid hydrodynamics, quantum sensing and matter-wave interferometry. The characterization of the rotational properties of such superfluids is important, but can presently only be performed by using optical absorption imaging, which completely destroys the condensate. Recent studies have proposed coupling the ring BEC to optical cavity modes carrying orbital angular momentum to make minimally destructive measurements of the condensate rotation. The sensitivity of these proposals, however, is bounded below by the standard quantum limit set by the combination of laser shot noise and radiation pressure noise. In this work, we provide a theoretical framework that exploits the fact that the interaction between the scattered modes of the condensate and the light reduces to effective optomechanical equations of motion. We present a detailed theoretical analysis to demonstrate that the use of squeezed light and backaction evasion techniques allows the angular momentum of the condensate to be sensed with noise well below the standard quantum limit. Our proposal is relevant to atomtronics, quantum sensing and quantum information.

quant-ph↗

FLIRT: Feedback Loop In-context Red Teaming

Warning: this paper contains content that may be inappropriate or offensive. As generative models become available for public use in various applications, testing and analyzing vulnerabilities of these models has become a priority. In this work, we propose an automatic red teaming framework that evaluates a given black-box model and exposes its vulnerabilities against unsafe and inappropriate content generation. Our framework uses in-context learning in a feedback loop to red team models and trigger them into unsafe content generation. In particular, taking text-to-image models as target models, we explore different feedback mechanisms to automatically learn effective and diverse adversarial prompts. Our experiments demonstrate that even with enhanced safety features, Stable Diffusion (SD) models are vulnerable to our adversarial prompts, raising concerns on their robustness in practical uses. Furthermore, we demonstrate that the proposed framework is effective for red teaming text-to-text models.

cs.AI↗

Strategic Electric Distribution Network Sensing via Spectral Bandits

Despite their wide-scale deployment and ability to make accurate high-frequency voltage measurements, communication network limitations have largely precluded the use of smart meters for real-time monitoring purposes in electric distribution systems. Although smart meter communication networks have limited bandwidth available per meter, they also have the ability to dedicate higher bandwidth to varying subsets of meters. Using this capability to enable real-time monitoring from smart meters, this paper proposes an online bandwidth-constrained sensor sampling algorithm that takes advantage of the graphical structure inherent in the power flow equations. The key idea is to use a spectral bandit framework where the estimated parameters are the graph Fourier transform coefficients of the nodal voltages. The structure provided by this framework promotes a sampling policy that strategically accounts for electrical distance. Maxima of sub-Gaussian random variables model the policy rewards, which relaxes distributional assumptions common in prior work. The scheme is implemented on a synthetic electrical network to dynamically identify meters exposing violations of voltage magnitude limits, illustrating the effectiveness of the proposed method.

eess.SY↗

Self-Contradictory Reasoning Evaluation and Detection

In a plethora of recent work, large language models (LLMs) demonstrated impressive reasoning ability, but many proposed downstream reasoning tasks only focus on final answers. Two fundamental questions persist: 1) how consistent is the reasoning, and 2) can models detect unreliable reasoning? In this paper, we investigate self-contradictory (Self-Contra) reasoning, where the model reasoning does not support its answers. To answer 1), we define and assess the Self-Contra rate across three datasets and delve into finer-grained categories of Self-Contra reasoning. We find that LLMs often contradict themselves in reasoning tasks involving contextual information understanding or commonsense. The model may generate correct answers by taking shortcuts in reasoning or overlooking contextual evidence, leading to compromised reasoning. For 2), we task the state-of-the-art model GPT-4 with identifying Self-Contra reasoning and finer-grained fallacies. We find that finer-grained categories enhanced detection can improve GPT-4's ability to detect Self-Contra. However, it is only able to detect Self-Contra with a 52.2% F1 score, much lower compared to 66.7% for humans. Our results indicate that current LLMs lack the robustness necessary for reliable reasoning and we emphasize the urgent need for establishing best practices in comprehensive reasoning evaluations beyond pure performance-based metrics.

cs.CL↗

Terahertz emission from $α$-W/CoFe epitaxial spintronic emitters

We report efficient terahertz (THz) generation in epitaxial $α$-W/Co$_{60}$Fe$_{40}$ spintronic emitters. Two types of emitters have been investigated; epitaxial $α$-W$(110)$/Co$_{60}$Fe$_{40}(110)$ and $α$-W$(001)$/Co$_{60}$Fe$_{40}(001)$ deposited on single crystalline Al$_{2}$O$_{3}$($11\bar{2}0$) and MgO($001$) substrates, respectively. First principle calculations of the electronic band structure at the W$(001)$ surface reveal Dirac-type surface states, similar to that reported previously for the W$(110)$ surface. The generated THz radiation is about $10\%$ larger for $α$-W$(110)$/Co$_{60}$Fe$_{40}(110)$ grown on single crystalline Al$_{2}$O$_{3}$($11\bar{2}0$), which is explained by the fact that the $α$-W$(110)$/Co$_{60}$Fe$_{40}(110)$ interface for this emitter is more transparent to the spin current due to the presence of Ångstr\" om-scale interface intermixing at the W/CoFe interface. Our results also reveal that the generation of THz radiation is larger when pumping with the laser light from the substrate side, which is explained by a larger part of the laser light due to interference effects in the film stack being absorbed in the ferromagnetic Co$_{60}$Fe$_{40}$ layer in this measurement configuration.

physics.app-ph↗