Searcharxiv⌕ Search

arXiv subjects

Rahul Gupta

Publications and source records attributed to Rahul Gupta.

At least 37 records · Page 2Linked to original sources

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

Post-training Large Language Models requires diverse, high-quality data which is rare and costly to obtain, especially in low resource domains and for multi-turn conversations. Common solutions are crowdsourcing or synthetic generation, but both often yield low-quality or low-diversity data. We introduce Adversarial Arena for building high quality conversational datasets by framing data generation as an adversarial task: attackers create prompts, and defenders generate responses. This interactive competition between multiple teams naturally produces diverse and complex data. We validated this approach by conducting a competition with 10 academic teams from top US and European universities, each building attacker or defender bots. The competition, focused on safety alignment of LLMs in cybersecurity, generated 19,683 multi-turn conversations. Fine-tuning an open-source model on this dataset produced an 18.47% improvement in secure code generation on CyberSecEval-Instruct and 29.42% improvement on CyberSecEval-MITRE.

cs.AI↗

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an imperfect Reward Model (RM) can become a single point of failure when it fails to penalize unsafe behaviors. While existing red-teaming approaches primarily target policy-level weaknesses, they overlook what we term systemic weaknesses cases where both the core LLM and the RM fail in tandem. We present ARES, a framework that systematically discovers and mitigates such dual vulnerabilities. ARES employs a ``Safety Mentor'' that dynamically composes semantically coherent adversarial prompts by combining structured component types (topics, personas, tactics, goals) and generates corresponding malicious and safe responses. This dual-targeting approach exposes weaknesses in both the core LLM and the RM simultaneously. Using the vulnerabilities gained, ARES implements a two-stage repair process: first fine-tuning the RM to better detect harmful content, then leveraging the improved RM to optimize the core model. Experiments across multiple adversarial safety benchmarks demonstrate that ARES substantially enhances safety robustness while preserving model capabilities, establishing a new paradigm for comprehensive RLHF safety alignment.

cs.AI↗

Optimally Controlled Storage of a Qubit in an Inhomogeneous Spin Ensemble

The storage of quantum information in spin-ensembles is limited by practically unavoidable inhomogeneous broadening, and the macroscopic number of spins in such an ensemble makes the design of control solutions to increase the coherence time a challenging task. Together with a concurrently developed Krylov theory that allows us to treat the control problem efficiently, we design optimal cavity modulation for such spin ensembles that achieve an order of magnitude enhancement in qubit lifetime compared to the losses due to inhomogeneity and cavity decay.

quant-ph↗

Quantum information spreading in inhomogeneous spin ensembles

We present a Krylov space based theoretical framework for modeling inhomogeneous spin ensembles with arbitrary distributions of spin frequencies and couplings. The framework is then used to asymptotically large spin ensemble. In the single-excitation subspace, the Krylov construction allows for to derive exact expressions for the Lieb-Robinson velocity and quantum speed limit, and figure of merit such as Krylov complexity. Our work reveals a strong dependence of the speed of information flow on the statistical distribution of resonance frequencies in the spin ensemble with immediate implications for the design of components for quantum technologies, realized for example with nitrogen vacancy centers, nuclear spins or ultracold atoms.

quant-ph↗

Motivic Cohomology and K-groups of varieties over higher local fields

For quasi-projective varieties over a higher local field $k_N$, we prove that its $K$-groups, above a suitable degree, are divisible-by-finite. We also prove the finiteness of the prime-to-$p$ torsion subgroup of certain higher Chow groups for smooth projective varieties over such fields, where $p$ denotes the final residue characteristic of $k_N$. As an application, we show that the kernel of the tame reciprocity map is uniquely $p'$-divisible. A key ingredient in achieving these results is the finiteness of étale cohomology groups over such fields.

math.AG↗

Large Language Model-driven Analysis of General Coordinates Network (GCN) Circulars

The General Coordinates Network (GCN) is NASA's time-domain and multimessenger alert system. GCN distributes two data products: automated "Notices" and human-generated "Circulars" that report the observations of high-energy and multimessenger astronomical transients. The flexible and nonstructured format of GCN Circulars, comprising more than 40,500 Circulars accumulated over three decades, makes it challenging to manually extract observational information, such as redshift or observed wave bands. In this work, we employ large language models (LLMs) to facilitate the automated parsing of transient reports. We develop a neural topic modeling pipeline with open-source tools for the automatic clustering and summarization of astrophysical topics in the Circulars archive. Using neural topic modeling and contrastive fine-tuning, we classify Circulars based on their observation wave bands and messengers. Additionally, we separate gravitational-wave event clusters and their electromagnetic counterparts from the Circulars archive. Finally, using the open-source Mistral model, we implement a system to automatically extract gamma-ray burst (GRB) redshift information from the Circulars archive, without the need for any training. Evaluation against the manually curated Neil Gehrels Swift Observatory GRB table shows that our simple system, with the help of prompt-tuning, output parsing, and retrieval augmented generation (RAG), can achieve an accuracy of 97.2% for redshift-containing Circulars. Our neural search-enhanced RAG pipeline accurately retrieved 96.8% of redshift Circulars from the manually curated archive. Our study demonstrates the potential of LLMs to automate and enhance astronomical text mining and provides a foundational work for future advances in transient alert analysis.

astro-ph.HE↗

Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks

Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Defending against novel jailbreaks represents a critical challenge in AI safety. Adversarial training -- designed to make models robust against worst-case perturbations -- has been the dominant paradigm for adversarial robustness. However, due to optimization challenges and difficulties in defining realistic threat models, adversarial training methods often fail on newly developed jailbreaks in practice. This paper proposes a new paradigm for improving robustness against unseen jailbreaks, centered on the Adversarial Déjà Vu hypothesis: novel jailbreaks are not fundamentally new, but largely recombinations of adversarial skills from previous attacks. We study this hypothesis through a large-scale analysis of 32 attack papers published over two years. Using an automated pipeline, we extract and compress adversarial skills into a sparse dictionary of primitives, with LLMs generating human-readable descriptions. Our analysis reveals that unseen attacks can be effectively explained as sparse compositions of earlier skills, with explanatory power increasing monotonically as skill coverage grows. Guided by this insight, we introduce Adversarial Skill Compositional Training (ASCoT), which trains on diverse compositions of skill primitives rather than isolated attack instances. ASCoT substantially improves robustness to unseen attacks, including multi-turn jailbreaks, while maintaining low over-refusal rates. We also demonstrate that expanding adversarial skill coverage, not just data scale, is key to defending against novel attacks. \textcolor{red}{\textbf{Warning: This paper contains content that may be harmful or offensive in nature.

cs.LG↗

Episode-wise spectro-polarimetry of GRB 220107A: Testing the hypothesis of evolving radiation mechanisms

We investigate the spectro-polarimetric properties of the long-duration GRB~220107A, which exhibited two distinct emission episodes separated by a 40 s quiescent gap, to test whether such multi-episode bursts show evidence for evolution in their underlying radiation mechanisms. We analyzed prompt emission data from AstroSat/CZTI, Fermi/GBM, and Konus-Wind, performing spectro-polarimetric analysis for each emission episode. The time-integrated polarization analysis shows no significant detection (PF$ < 38 \%$, $2σ$). Time-resolved analysis reveals clear spectral evolution between the two episodes, with episode 1 exhibiting a hard low-energy photon index and episode 2 showing substantial spectral softening ($α\sim -0.72$). Regarding polarization: Episode 1 shows a low polarization upper limit (< 52\%), consistent with expectations for photospheric emission dominated by quasi-thermal Comptonization in a baryon-rich outflow. Episode 2 also shows overall low polarization (PF$ < 55 \%$, $2σ$), though sliding-window analysis yields a marginally elevated signal (PF$= 70 \pm 30\%$, BF = 2.8) between T0+76 to T0+88 s. The robust spectral softening between episodes could arise from sub-photospheric dissipation, optically thin synchrotron radiation in small-scale magnetic fields, or if the tentative polarization enhancement proves intrinsic, it would favor synchrotron emission in large-scale ordered magnetic fields. The spectral evolution of GRB 220107A, combined with our polarimetric constraints, demonstrates the diagnostic potential of time-resolved spectro-polarimetry for constraining GRB prompt emission physics. We present GRB 220107A as a test case illustrating how future higher sensitivity observations could discriminate between competing emission models for multi-episode bursts. Our results emphasize both the promise and current limitations of prompt phase polarimetry.

astro-ph.HE↗

How Catastrophic is Your LLM? Certifying Risk in Conversation

Large Language Models (LLMs) can produce catastrophic responses in conversational settings that pose serious risks to public safety and security. Existing evaluations often fail to fully reveal these vulnerabilities because they rely on fixed attack prompt sequences, lack statistical guarantees, and do not scale to the vast space of multi-turn conversations. In this work, we propose C$^3$LLM, a novel, principled statistical Certification framework for Catastrophic risks in multi-turn Conversation for LLMs that bounds the probability of an LLM generating catastrophic responses under multi-turn conversation distributions with statistical guarantees. We model multi-turn conversations as probability distributions over query sequences, represented by a Markov process on a query graph whose edges encode semantic similarity to capture realistic conversational flow, and quantify catastrophic risks using confidence intervals. We define several inexpensive and practical distributions--random node, graph path, and adaptive with rejection. Our results demonstrate that these distributions can reveal substantial catastrophic risks in frontier models, with certified lower bounds as high as 70% for the worst model, highlighting the urgent need for improved safety training strategies in frontier LLMs.

cs.AI↗

LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions

Deception is a pervasive feature of human communication and an emerging concern in large language models (LLMs). While recent studies document instances of LLM deception, most evaluations remain confined to single-turn prompts and fail to capture the long-horizon interactions in which deceptive strategies typically unfold. We introduce a new simulation framework, LH-Deception, for a systematic, empirical quantification of deception in LLMs under extended sequences of interdependent tasks and dynamic contextual pressures. LH-Deception is designed as a multi-agent system: a performer agent tasked with completing tasks and a supervisor agent that evaluates progress, provides feedback, and maintains evolving states of trust. An independent deception auditor then reviews full trajectories to identify when and how deception occurs. We conduct extensive experiments across 11 frontier models, spanning both closed-source and open-source systems, and find that deception is model-dependent, increases with event pressure, and consistently erodes supervisor trust. Qualitative analyses further reveal emergent, long-horizon phenomena, such as ``chains of deception", which are invisible to static, single-turn evaluations. Our findings provide a foundation for evaluating future LLMs in real-world, trust-sensitive contexts.

cs.CL↗

Evaluating Nova 2.0 Lite model under Amazon's Frontier Model Safety Framework

Amazon published its Frontier Model Safety Framework (FMSF) as part of the Paris AI summit, following which we presented a report on Amazon's Premier model. In this report, we present an evaluation of Nova 2.0 Lite. Nova 2.0 Lite was made generally available from amongst the Nova 2.0 series and is one of its most capable reasoning models. The model processes text, images, and video with a context length of up to 1M tokens, enabling analysis of large codebases, documents, and videos in a single prompt. We present a comprehensive evaluation of Nova 2.0 Lite's critical risk profile under the FMSF. Evaluations target three high-risk domains-Chemical, Biological, Radiological and Nuclear (CBRN), Offensive Cyber Operations, and Automated AI R&D-and combine automated benchmarks, expert red-teaming, and uplift studies to determine whether the model exceeds release thresholds. We summarize our methodology and report core findings. We will continue to enhance our safety evaluation and mitigation pipelines as new risks and capabilities associated with frontier models are identified.

cs.CR↗

Tame class field theory over local fields

For a quasi-projective scheme $X$ admitting a smooth compactification over a local field of residue characteristic $p > 0$, we construct a continuous reciprocity homomorphism from a tame class group to the abelian tame etale fundamental group of $X$. We describe the prime-to-$p$ parts of its kernel and cokernel. This generalizes the higher dimensional unramified class field theory over local fields by Jannsen-Saito and Forre. We also prove a finiteness theorem for the geometric part of the abelian tame etale fundamental group, generalizing the results of Grothendieck and Yoshida for the unramified fundamental group.

math.AG↗

Embedded AI Companion System on Edge Devices

Computational resource constraints on edge devices make it difficult to develop a fully embedded AI companion system with a satisfactory user experience. AI companion and memory systems detailed in existing literature cannot be directly used in such an environment due to lack of compute resources and latency concerns. In this paper, we propose a memory paradigm that alternates between active and inactive phases: during phases of user activity, the system performs low-latency, real-time dialog using lightweight retrieval over existing memories and context; whereas during phases of user inactivity, it conducts more computationally intensive extraction, consolidation, and maintenance of memories across full conversation sessions. This design minimizes latency while maintaining long-term personalization under the tight constraints of embedded hardware. We also introduce an AI Companion benchmark designed to holistically evaluate the AI Companion across both its conversational quality and memory capabilities. In our experiments, we found that our system (using a very weak model: Qwen2.5-7B-Instruct quantized int4) outperforms the equivalent raw LLM without memory across most metrics, and performs comparably to GPT-3.5 with 16k context window.

cs.AI↗

Cavity-Driven Multispectral Gain for High-Sensitivity NV Center Magnetometers

We report a cavity-enabled solid-state magnetometer based on an NV ensemble coupled with a dielectric cavity, achieving 12 pT/$\sqrt{\rm{Hz}}$ sensitivity and a nearly threefold gain from multispectral features. The features originate from cavity-induced splitting of the NV hyperfine levels and leverages robust quantum coherence in the doubly dressed states of the system to achieve high sensitivity. We project simulated near-term sensitivities approaching 100 fT/$\sqrt{\rm{Hz}}$, close to the Johnson-Nyquist limit. Our results establish frequency multiplexing as a new operational paradigm, offering a robust and scalable quantum resource for metrology under ambient conditions.

quant-ph↗

Macroscopic entanglement between localized domain walls inside a cavity

We present a scheme for generating stable and tunable entanglement between two localized Bloch domain walls in nanomagnetic strips kept inside a chiral optical cavity. The entanglement is mediated by the effective optomechanical interaction between the cavity photons and the two macroscopic, collective modes of the pinned domain walls. By controlling the pinning potential and optical driving frequency, the robust, steady-state entanglement between the two macroscopic domain walls can survive beyond the typical milli-Kelvin temperature range.

quant-ph↗

S&P 500 Stock's Movement Prediction using CNN

This paper is about predicting the movement of stock consist of S&P 500 index. Historically there are many approaches have been tried using various methods to predict the stock movement and being used in the market currently for algorithm trading and alpha generating systems using traditional mathematical approaches [1, 2]. The success of artificial neural network recently created a lot of interest and paved the way to enable prediction using cutting-edge research in the machine learning and deep learning. Some of these papers have done a great job in implementing and explaining benefits of these new technologies. Although most these papers do not go into the complexity of the financial data and mostly utilize single dimension data, still most of these papers were successful in creating the ground for future research in this comparatively new phenomenon. In this paper, I am trying to use multivariate raw data including stock split/dividend events (as-is) present in real-world market data instead of engineered financial data. Convolution Neural Network (CNN), the best-known tool so far for image classification, is used on the multi-dimensional stock numbers taken from the market mimicking them as a vector of historical data matrices (read images) and the model achieves promising results. The predictions can be made stock by stock, i.e., a single stock, sector-wise or for the portfolio of stocks.

cs.CV↗

Itinerant Orbital Hall Effect Mechanism Leading to Large Negative Orbital Torques from Light Metal Vanadium

The orbital Hall effect (OHE) has attracted significant attention for developing energy-efficient electronic devices. However, utilizing it in fast, low-power devices requires an enhanced understanding of underlying extrinsic and intrinsic contributions to OHE at timescales ranging from quasi-static to picoseconds. Here, we investigate OHE in light metal vanadium (V) using a combination of selected measurement schemes, spanning the full frequency range. We observe a negative damping-like torque efficiency from V, opposite to conventional theoretical predictions, with a magnitude that depends on the adjacent ferromagnet, a dependence that indicates orbital effects. These results, with consistent torque efficiencies across all frequencies, corroborate a negative and intrinsic OHE in V with a large effective orbital Hall conductivity of $-(1.44 \pm 0.34)\,\frac{\hbar}{2e}\,\times 10^{5}\,Ω^{-1}\,\mathrm{m}^{-1}$ and a long orbital diffusion length of $(15.0 \pm 2.5)\,\mathrm{nm}$. To explain the observed OHE, we develop a theoretical model incorporating both local and itinerant circulation contributions to OHE. The model agrees excellently with the experimental results, demonstrating that itinerant contributions are essential for a complete physical understanding of intrinsic OHE. Our consistent experimental and theoretical data highlight the importance of itinerant contributions governing the fundamental understanding of intrinsic OHE and the large effects found open pathways for energy-efficient orbitronic devices.

cond-mat.mes-hall↗

Extremely luminous optical afterglow of an energetic gamma-ray burst GRB 230204B

Robotic telescope networks play an important role in capturing early and bright optical afterglows, providing critical insights into the energetics and emission mechanisms of GRBs. In this study, we analyze GRB 230204B, an exceptionally energetic and multi-pulsed long GRB, detected by the Fermi GBM and MAXI detectors, with an isotropic equivalent gamma-ray energy exceeding 10$^{54}$ erg. Time-resolved spectral analysis reveals a transition in the prompt emission from hard (sub-photospheric dominated) spectra during early pulses to softer (synchrotron radiation dominated) spectra in later pulses, indicative of a hybrid jet composition. We report the discovery and characterization of the optical afterglow using the MASTER and BOOTES robotic telescope networks, which enabled rapid follow-up observations starting at $\sim$1.3 ks post-burst. The optical luminosity at this time was exceptionally high, surpassing that of many other optically bright GRBs, such as GRB 990123, GRB 080319B, etc. This places the burst among the most luminous optical GRBs observed to date. Long-term radio observations extending to 335 days post-burst were conducted with the ATCA. Multi-wavelength modeling was conducted using an external ISM forward-shock top-hat jet model with \sw{afterglowpy}. The results reveal a narrow and highly collimated jet with a circumburst density of $n_{0} \sim$ 28.12 cm$^{-3}$, kinetic energy $E_{\rm K} \sim$ 4.18 $\times 10^{55}$ erg, and a relatively low value of $ε_{B}$ = 2.14 $\times 10^{-6}$, indicating shock-compression of magnetic field in the surrounding interstellar medium. We constrained a low radiative efficiency of $\sim$ 4.3 \%. This study highlights the indispensable contribution of robotic networks to early afterglow observations and advances our understanding of GRB 230204B unique characteristics and underlying jet physics.

astro-ph.HE↗