SearcharxivSearch

arXiv subjects

Shreyans Jain

Publications and source records attributed to Shreyans Jain.

11 recordsLinked to original sources

Gotta Catch them all: the modes of Sycophancy

Large language models often align with users' beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies largely treat sycophancy as a single behavioral dimension that can be uniformly amplified or suppressed. We challenge this assumption by analyzing three hypothesized modes of sycophancy across 948 social pressure situations. Although the modes produce highly similar outputs, with a text-only classifier achieving just 57.8 percent accuracy, their internal representations are perfectly linearly separable from layer 14 onward. We further find the modes emerge at different processing stages, rely on distinct attention circuitry, and fire strongest on different inputs. These results show that sycophancy is not a monolithic tendency, but a structured family of representationally and computationally distinct modes, motivating more precise measurement and intervention.

cs.CL

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We propose a framework to measure consistent behavioral tendencies using situational judgment tests (SJTs), multidimensional item response theory (MIRT), and structured synthetic personas, treating responses as observations of latent behavioral variables. Across large-scale SJT and persona datasets, we find that persona-conditioned behaviors are stable across runs, latent trait scores predict external benchmarks (e.g., TruthfulQA, EmoBench), and MIRT reveals consistent latent structure. We validate these results through human annotation, benchmark evaluation, and internal consistency analyses. We interpret these traits not as human personality, but as stable behavioral tendencies expressed across contexts. Our results show that scenario-based psychometric evaluation provides a more reliable alternative to classical self-report approaches for assessing LLM behavior, and we release datasets to support further study.

cs.AI

Unlearning in Diffusion models under Data Constraints: A Variational Inference Approach

For a responsible and safe deployment of diffusion models in various domains, regulating the generated outputs from these models is desirable because such models could generate undesired, violent, and obscene outputs. To tackle this problem, recent works use machine unlearning methodology to forget training data points containing these undesired features from pre-trained generative models. However, these methods proved to be ineffective in data-constrained settings where the whole training dataset is inaccessible. Thus, the principal objective of this work is to propose a machine unlearning methodology that can prevent the generation of outputs containing undesired features from a pre-trained diffusion model in such a data-constrained setting. Our proposed method, termed as Variational Diffusion Unlearning (VDU), is a computationally efficient method that only requires access to a subset of training data containing undesired features. Our approach is inspired by the variational inference framework with the objective of minimizing a loss function consisting of two terms: plasticity inducer and stability regularizer. Plasticity inducer reduces the log-likelihood of the undesired training data points, while the stability regularizer, essential for preventing loss of image generation quality, regularizes the model in parameter space. We validate the effectiveness of our method through comprehensive experiments for both class unlearning and feature unlearning. For class unlearning, we unlearn some user-identified classes from MNIST, CIFAR-10, and tinyImageNet datasets from a pre-trained unconditional denoising diffusion probabilistic model (DDPM). Similarly, for feature unlearning, we unlearn the generation of certain high-level features from a pre-trained Stable Diffusion model trained on LAION-5B dataset.

cs.LG

Non-linear cooling and control of a mechanical quantum harmonic oscillator

Non-linearities are a key feature allowing non-classical control of quantum harmonic oscillators. However, when non-linearities are strong, designing protocols for control is often difficult, placing a barrier to exploiting these properties fully. Here, using a single trapped-ion oscillator operated in the strongly non-linear regime of the atom-light interaction, we show how to generate localized multi (2, 3, 4, and 5)-component Schr\"odinger's cat manifolds using a novel form of non-linear reservoir engineering. We then specifically select Hamiltonians which allow us to perform measurements on these state manifolds. To our knowledge, our work is the first experimental use of such high order non-linear processes for control of non-classical states of a quantum harmonic oscillator, opening up a new toolbox which can be applied to bosonic quantum error correction, computation, and sensing.

quant-ph

Sycophancy as compositions of Atomic Psychometric Traits

Sycophancy is a key behavioral risk in LLMs, yet is often treated as an isolated failure mode that occurs via a single causal mechanism. We instead propose modeling it as geometric and causal compositions of psychometric traits such as emotionality, openness, and agreeableness - similar to factor decomposition in psychometrics. Using Contrastive Activation Addition (CAA), we map activation directions to these factors and study how different combinations may give rise to sycophancy (e.g., high extraversion combined with low conscientiousness). This perspective allows for interpretable and compositional vector-based interventions like addition, subtraction and projection; that may be used to mitigate safety-critical behaviors in LLMs.

cs.AI

Beyond Linear Steering: Unified Multi-Attribute Control for Language Models

Controlling multiple behavioral attributes in large language models (LLMs) at inference time is a challenging problem due to interference between attributes and the limitations of linear steering methods, which assume additive behavior in activation space and require per-attribute tuning. We introduce K-Steering, a unified and flexible approach that trains a single non-linear multi-label classifier on hidden activations and computes intervention directions via gradients at inference time. This avoids linearity assumptions, removes the need for storing and tuning separate attribute vectors, and allows dynamic composition of behaviors without retraining. To evaluate our method, we propose two new benchmarks, ToneBank and DebateMix, targeting compositional behavioral control. Empirical results across 3 model families, validated by both activation-based classifiers and LLM-based judges, demonstrate that K-Steering outperforms strong baselines in accurately steering multiple behaviors.

cs.LG

A 3-dimensional scanning trapped-ion probe

Single-atom quantum sensors offer high spatial resolution and high sensitivity to electric and magnetic fields. Among them, trapped ions offer exceptional performance in sensing electric fields, which has been used in particular to probe these in the proximity of metallic surfaces. However, the flexibility of previous work was limited by the use of radio-frequency trapping fields, which has restricted spatial scanning to linear translations, and calls into question whether observed phenomena are connected to the presence of the radio-frequency fields. Here, using a Penning trap instead, we demonstrate a single ion probe which offers three-dimensional position scanning at distances between $50$ $\mu\mathrm{m}$ and $450$ $\mu\mathrm{m}$ from a metallic surface and above a $200\times200$ $\mu\mathrm{m}^{2}$ area, allowing us to reconstruct static and time-varying electric as well as magnetic fields. We use this to map charge distributions on the metallic surface and noise stemming from it. The methods demonstrated here allow similar probing to be carried out on samples with a variety of materials, surface constitutions and geometries, providing a new tool for surface science.

quant-ph

WavShadow: Wavelet Based Shadow Segmentation and Removal

Shadow removal and segmentation remain challenging tasks in computer vision, particularly in complex real world scenarios. This study presents a novel approach that enhances the ShadowFormer model by incorporating Masked Autoencoder (MAE) priors and Fast Fourier Convolution (FFC) blocks, leading to significantly faster convergence and improved performance. We introduce key innovations: (1) integration of MAE priors trained on Places2 dataset for better context understanding, (2) adoption of Haar wavelet features for enhanced edge detection and multiscale analysis, and (3) implementation of a modified SAM Adapter for robust shadow segmentation. Extensive experiments on the challenging DESOBA dataset demonstrate that our approach achieves state of the art results, with notable improvements in both convergence speed and shadow removal quality.

cs.CV

Penning micro-trap for quantum computing

Trapped ions in radio-frequency traps are among the leading approaches for realizing quantum computers, due to high-fidelity quantum gates and long coherence times. However, the use of radio-frequencies presents a number of challenges to scaling, including requiring compatibility of chips with high voltages, managing power dissipation and restricting transport and placement of ions. By replacing the radio-frequency field with a 3 T magnetic field, we here realize a micro-fabricated Penning ion trap which removes these restrictions. We demonstrate full quantum control of an ion in this setting, as well as the ability to transport the ion arbitrarily in the trapping plane above the chip. This unique feature of the Penning micro-trap approach opens up a modification of the Quantum CCD architecture with improved connectivity and flexibility, facilitating the realization of large-scale trapped-ion quantum computing, quantum simulation and quantum sensing.

quant-ph

Engineering generalized Gibbs ensembles with trapped ions

The concept of generalized Gibbs ensembles (GGEs) has been introduced to describe steady states of integrable models. Recent advances show that GGEs can also be stabilized in nearly integrable quantum systems when driven by external fields and open. Here, we present a weakly dissipative dynamics that drives towards a steady-state GGE and is realistic to implement in systems of trapped ions. We outline the engineering of the desired dissipation by a combination of couplings which can be realized with ion-trap setups and discuss the experimental observables needed to detect a deviation from a thermal state. We present a novel mixed-species motional mode engineering technique in an array of micro-traps and demonstrate the possibility to use sympathetic cooling to construct many-body dissipators. Our work provides a blueprint for experimental observation of GGEs in open systems and opens a new avenue for quantum simulation of driven-dissipative quantum many-body problems.

quant-ph

Scalable arrays of micro-Penning traps for quantum computing and simulation

We propose the use of 2-dimensional Penning trap arrays as a scalable platform for quantum simulation and quantum computing with trapped atomic ions. This approach involves placing arrays of micro-structured electrodes defining static electric quadrupole sites in a magnetic field, with single ions trapped at each site and coupled to neighbors via the Coulomb interaction. We solve for the normal modes of ion motion in such arrays, and derive a generalized multi-ion invariance theorem for stable motion even in the presence of trap imperfections. We use these techniques to investigate the feasibility of quantum simulation and quantum computation in fixed ion lattices. In homogeneous arrays, we show that sufficiently dense arrays are achievable, with axial, magnetron and cyclotron motions exhibiting inter-ion dipolar coupling with rates significantly higher than expected decoherence. With the addition of laser fields these can realize tunable-range interacting spin Hamiltonians. We also show how local control of potentials allows isolation of small numbers of ions in a fixed array and can be used to implement high fidelity gates. The use of static trapping fields means that our approach is not limited by power requirements as system size increases, removing a major challenge for scaling which is present in standard radio-frequency traps. Thus the architecture and methods provided here appear to open a path for trapped-ion quantum computing to reach fault-tolerant scale devices.

quant-ph