SearcharxivSearch

arXiv subjects

Ayan Datta

Publications and source records attributed to Ayan Datta.

12 recordsLinked to original sources

Efficient Safety Benchmarking via Item Response Theory

Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly problematic for adversarial, highly heterogeneous safety items. Applied in full to modern benchmark suites, the current evaluation procedures would require on the order of $10^5$ responses, most of which provide little ranking signal. We analyze a suite of widely used safety benchmarks and make three contributions toward more efficient safety evaluation. First, we show that Item Response Theory (IRT) recovers interpretable structure on safety benchmarks, with ability estimates resolving differences among models that cluster at the ceiling of raw safety metrics. Second, we show that adaptive item selection, which dynamically chooses informative items for each model based on its responses, approximates full-benchmark rankings while reducing evaluation cost by at least 80% on benchmarks where Spearman's $\rho >$90% with full-benchmark is attainable, and by up to 99.9% on AIR-Bench 2024. Third, we introduce a practical procedure for extracting a fixed, informative subset of items reusable across models, providing an alternative to adaptive selection with savings of up to 99.8% on AIR-Bench 2024. Together, these results establish that psychometric methods enable benchmark-aware reductions in evaluation costs across the safety evaluation pipeline.

cs.CY

Large Language Models Decide Early and Explain Later

Large Language Models often achieve strong performance by generating long intermediate chain-of-thought reasoning. However, it remains unclear when a model's final answer is actually determined during generation. If the answer is already fixed at an intermediate stage, subsequent reasoning tokens may constitute post-decision explanation, increasing inference cost and latency without improving correctness. We study the evolution of predicted answers over reasoning steps using forced answer completion, which elicits the model's intermediate predictions at partial reasoning prefixes. Focusing on Qwen3-4B and averaging results across all datasets considered, we find that predicted answers change in only 32% of queries. Moreover, once the final answer switch occurs, the model generates an average of 760 additional reasoning tokens per query, accounting for a substantial fraction of the total reasoning budget. Motivated by these findings, we investigate early stopping strategies that halt generation once the answer has stabilized. We show that simple heuristics, including probe-based stopping, can reduce reasoning token usage by 500 tokens per query while incurring only a 2% drop in accuracy. Together, our results indicate that a large portion of chain-of-thought generation is redundant and can be reduced with minimal impact on performance.

cs.CL

From Early Encoding to Late Suppression: Interpreting LLMs on Character Counting Tasks

Large language models (LLMs) exhibit failures on elementary symbolic tasks such as character counting in a word, despite excelling on complex benchmarks. Although this limitation has been noted, the internal reasons remain unclear. We use character counting (e.g., "How many p's are in apple?") as a minimal, controlled probe that isolates token-level reasoning from higher-level confounds. Using this setting, we uncover a consistent phenomenon across modern architectures, including LLaMA, Qwen, and Gemma: models often compute the correct answer internally yet fail to express it at the output layer. Through mechanistic analysis combining probing classifiers, activation patching, logit lens analysis, and attention head tracing, we show that character-level information is encoded in early and mid-layer representations. However, this information is attenuated by a small set of components in later layers, especially the penultimate and final layer MLP. We identify these components as negative circuits: subnetworks that downweight correct signals in favor of higher-probability but incorrect outputs. Our results lead to two contributions. First, we show that symbolic reasoning failures in LLMs are not due to missing representations or insufficient scale, but arise from structured interference within the model's computation graph. This explains why such errors persist and can worsen under scaling and instruction tuning. Second, we provide evidence that LLM forward passes implement a form of competitive decoding, in which correct and incorrect hypotheses coexist and are dynamically reweighted, with final outputs determined by suppression as much as by amplification. These findings carry implications for interpretability and robustness: simple symbolic reasoning exposes weaknesses in modern LLMs, underscoring need for design strategies that ensure information is encoded and reliably used.

cs.CL

SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection

This document contains the details of the authors' submission to the proceedings of SemEval 2024's Task 8: Multigenerator, Multidomain, and Multilingual Black-Box Machine-Generated Text Detection Subtask A (monolingual) and B. Detection of machine-generated text is becoming an increasingly important task, with the advent of large language models (LLMs). In this paper, we lay out how using weighted averages of RoBERTa layers lets us capture information about text that is relevant to machine-generated text detection.

cs.CL

Pressure induced insulator-to-metal transition in few-layer FePS$_3$ at 1.5 GPa

In two-dimensional (2D) van der Waals (vdW) layered materials the application of pressure often induces a giant lattice collapse, which can subsequently drive an associated Mott transition. Here, we investigate room-temperature layer-dependent insulator-metal transition (IMT) and probable spin-crossover (SCO) in vdW magnet, FePS$_3$, under high-pressure using micro-Raman scattering. Experimentally obtained spectra, in agreement with the computed Raman modes, indicate evidence of IMT of FePS$_3$ started with a thickness-dependent critical pressure ($P_c$) which reduces to 1.5 GPa in trilayer flakes compared to 10.8 GPa for the bulk counterpart. Using a phenomenological model, we argue that strong structural anisotropy in few-layer flakes enhances the in-plane strain under applied pressure and is, therefore, ultimately responsible for reducing the critical pressure for the IMT with decreasing layer numbers. Reduction of the critical pressure for phase transition in vdW magnets to 1-2 GPa marks the possibility of using intercalated few-layers in the field-effect transistor device architecture, and thereby, avoiding the conventional use of the diamond anvil cell (DAC).

cond-mat.mes-hall

Ordered Short-Range Ripple Effects in Structures of Silicenes: Role of Puckering in the Aromatic Rings

Structural and electronic properties of the all-Si analogue of graphene, silicene have elucidated through DFT calculations. Silicene differs considerably from graphene in being `chair-type' puckered in each 6-membered ring which leads to ordered ripples across the surface. Binding energies suggest stability for such rippled silicenes and are predicted to behave as a finite gap semi-conductor with electron-hole symmetry quenched. Inter-layer coupling between the silicenes is suggested as the mechanism for the formation of the bulk-Si in its only known diamond form.

cond-mat.mtrl-sci

New examples of metalloaromatic Al-clusters: (Al4M4)Fe(CO)3 (M=Li, Na and K) and (Al4M4)2Ni: Rationalization for possible synthesis

Ab-initio calculations reveal that all-metal antiaromatic molecules like Al4M4 (M=Li, Na and K) can be stabilized in half-sandwich complex: (Al4M4)Fe(CO)3 and full-sandwich complexes of the type: (Al4M4)2Ni. The formation of the full-sandwich complex [(Al4M4)2Ni] from its organometallic precursor depends on the stability of the organic-inorganic hybrid (C4H4) Ni (Al4Li4).

physics.chem-ph

Rationalization of pi(sigma) anti(aromaticity) in all metal molecular clusters

A sigma-pi separation analysis of the energies in Al4Li4 reveals that the system is more pi-antiaromatic than the sigma-aromaticity in it. This is true also for C4H4 and Ga4Li4. Unlike C4H4 that has a very large component of pi-antiaromaticity, for these all-metal clusters, these energy scales are comparable though pi-antiaromaticity is the major driving force for the distortion of the these molecules from the square (sigma-aromatic) structure to the rectangular (pi-antiaromatic) architecture. For the dianion Al4Li42-, the sigma-equalization prevails over the pi-distortion in Al4Li4 and for the dication Al4Li42+, pi-equalization is the driving force for the square symmetric structure.

physics.chem-ph

Proton-Pump Mechanism in Retinal Schiff Base: On the molecular structure of the M-state

Theoretical characterizations of the various intermediates in the proton pump cycle of the retinal Schiff base in the Halobacterium salinarium have been performed. Contrary to the general belief over the years that the most stable intermediate, the M-state, is a non-protonated cis-isomer, we find that the M-state is a polarized cis-isomer stabilized due to interactions of the dissociating proton with the pi-electrons. The role of proton in the pump cycle is found to be profound leading to the stabilization or in certain cases destabilization of the intermediates. We propose the chemical structure of the M-state for the first time.

physics.chem-ph

Long range electron transfer across $π$-conjugated systems: role of electron correlations

We consider a prototype polyene chain: donor-$π$(bridge)-acceptor. The distance between the donor and the acceptor is varied by increasing the number of bridged atoms and rate of electron transfer, k$_{et}$ is studied for a series of different donors, D=NH$_2$, SH, OH, and a fixed acceptor, A=NO$_2$. We observe a large k$_{et}$ even at a very large D-A separation of $\sim$ 45 $Å$, unexpected from the well-known and standard theories like the Forster theory. Such a long range electron transfer is due to the strong electron-electron interactions in the bridged orbitals that result in deconfinment of electrons in donor orbitals. Calculations at various levels: semi-empirical and many-body exact, have been performed to accurately account for such correlations.

cond-mat.mes-hall

First examples of stable transition metal complexes of an all-metal antiaromatic molecule (Al$_4$Li$_4$)

We propose new methodologies for stabilizing all-metal antiaromatic clusters like: Al$_{4}$Li$_{4}$. We demonstrate that these all-metal species can be stabilized by complexation with 3d-transition metals very similar to its organic counterpart, C$_4$H$_4$. Complexation to transition metal ions reduce the frontier orbital energies and introduces aromaticity. We consider a series of such complexes [$η$$^4$(Al$_4$Li$_4$)-Fe(CO)$_3$, $η$$^2$$σ$$^2$(Al$_4$Li$_4$)-Ni and (Al$_{4}$Li$_{4}$)$_{2}$Ni] and make a comparison between the all-metal species and the organometallic compounds to prove conclusively our theory. Fragmentation energy analysis as well as NICS support similar mechanism of complexation induced stability in these all-metal molecules.

cond-mat.str-el

Charge-transfer induced large nonlinear optical properties of small Al Clusters: Al4M4

We investigate the linear and nonlinear electric polarizabilities of small Al4M4 (M=Li, Na and K) clusters. Quantum chemical calculations reveal that these compounds exhibit an exceptionally high magnitude of linear and nonlinear optical (NLO) coefficients which are orders of magnitude higher than the conventional pi-conjugated systems of similar sizes. We attribute such phenomenal increase to non-centrosymmetricity incorporated in the systems by the alkali atoms surrounding the ring leading to charge transfer with small optical gap and low bond length alternation (BLA). Such a low magnitude of the BLA from a different origin, suggests the possibility that these clusters are aromatic in character and along with the large NLO coefficients they appear to be better candidates for next generation NLO fabrication devices.

cond-mat.mtrl-sci