SearcharxivSearch

arXiv subjects

Shashank Srivastava

Publications and source records attributed to Shashank Srivastava.

At least 19 recordsLinked to original sources

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.

cs.CL

The Story Shapes the Agent: Narrative Priors in LLM Behavior

Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter more than the assigned persona. We isolate this effect through structural isomorphism, constructing three text-based investigation games that share the same action space, stage progression, and resource constraints while varying only task narrative: disease investigation, IT troubleshooting, and murder mystery. Across 1,890 sessions spanning 3 models and 10 personas, we identify narrative priors: systematic action tendencies activated by a task's story framing, independent of its decision structure. Narrative priors explain 5-31x more behavioral variance than persona, are consistent across model architectures, and in two of three domains are negatively associated with task success. Persona effects that do transfer across narratives arise from behavioral anchors, persona descriptions whose language maps directly onto shared actions. Causal interventions confirm this: removing anchor words from a high-transfer persona reduces cross-narrative consistency by 95%. Our framework also generalizes to a held-out fourth narrative and yields a persona-selection method that improves cross-narrative transfer. These results suggest that LLM behavior that survives narrative changes should be grounded in concrete actions rather than abstract descriptions.

cs.CL

Unconventional Superconductivity in the Chiral Topological Semimetal Ag2Pd3S

Chiral crystals provide a unique setting where broken inversion symmetry, strong spin-orbit coupling, and electronic topology intertwine, yet superconductivity in intrinsically chiral materials remains rare. Here, we report unconventional superconductivity in the chiral topological semimetal Ag$_2$Pd$_3$S, an enantiomorphic analog of natural mineral coldwellite, crystallizing in the right-handed space group $P4_132$. Bulk superconductivity with a transition temperature $T_C = 1.1(2)$ K is confirmed by electrical resistivity, magnetization, and specific-heat measurements. Muon spin rotation and relaxation ($\mu$SR) experiments reveal a fully gapped superconducting state that spontaneously time-reversal symmetry (TRS) breaking establishing Ag$_2$Pd$_3$S as the first chiral topological semimetal superconductor exhibiting intrinsic TRS breaking. First-principles calculations uncover multiple multifold band crossings near the Fermi level, hosting Kramers-Weyl, double spin-1, and spin-3/2 quasiparticles with large topological charges. These unconventional fermions generate symmetry-protected topological surface states and underscore the nontrivial topology of the normal state. Symmetry analysis based on the Ginzburg-Landau theory suggests a loop-supercurrent-ordered superconducting state, yielding a full gap alongside spontaneous TRS breaking. The coexistence of TRS-breaking superconductivity and chiral multifold fermions identifies Ag$_2$Pd$_3$S as a platform for realizing intrinsic superconducting diode effects and chirality-induced spin selectivity, offering a transformative pathway toward dissipationless topological quantum technologies.

cond-mat.supr-con

The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoning

Existing deception datasets label completed outputs as honest or deceptive, treating deception as a property of the final response rather than a function of the model's reasoning trace. This obscures a more fundamental question: when does a language model become committed to deception? We introduce counterfactual localization: for each sentence prefix in a reasoning trace, we fix the prefix, resample continuations, and estimate the probability of a deceptive outcome. To scale this, we construct five environments (spanning strategic bluffing, maze guidance, financial advice, used-car sales, and offer negotiation) in which deception is never prompted but emerges from strategic incentives and labels follow mechanically from environment state rather than subjective human judgment. The resulting corpus localizes $\sim$1.46M sentences across four reasoning models, drawn from over 94.1M sampled continuations, 91.5B generated tokens, and over 100K scenarios. Sentence-level human evaluation confirms that detected commitment points correspond to interpretable shifts in decision state. Using this resource, we show that lexical cues for commitment prediction transfer poorly across environments, whereas attention-based transition features generalize out of distribution, suggesting that deceptive commitment is reflected in reusable changes in reasoning dynamics rather than surface form. We further identify compact attention-head sets (under 10% of heads) that, selected on one environment, causally suppress deceptive commitment across held-out environments. We release the corpus as a substrate for studying deception, and more broadly commitment, in language-model reasoning.

cs.CL

Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization

Recent work, using the Biasing Features metric, labels a CoT as unfaithful if it omits a prompt-injected hint that affected the prediction. We argue this metric adopts a narrow notion of faithfulness and confuses unfaithfulness with incompleteness, the lossy compression needed to turn distributed transformer computation into a linear natural language narrative. On multi-hop reasoning tasks with instruct-tuned and reasoning models, many CoTs flagged as unfaithful by Biasing Features are judged faithful by other metrics, exceeding 50% in some models. With a new faithful@k metric, we show that larger inference-time budgets greatly increase hint verbalization (up to 90% in some settings), suggesting much apparent unfaithfulness is due to tight token limits. Using Causal Mediation Analysis, we further show that even non-verbalized hints can causally mediate prediction changes through the CoT. We therefore caution against relying solely on hint-based evaluations and advocate a broader interpretability toolkit, including causal mediation and corruption-based metrics. We do not claim all CoTs are faithful, only that the absence of hint words alone does not prove unfaithfulness.

cs.CL

Discovery of Quasi One Dimensional Superconductivity in PtPb3Bi

Quasi one dimensional materials provide a compelling platform where reduced dimensionality stabilizes intertwined topological and superconducting phases. Here we report superconductivity in a new Bi based quasi 1D compound, PtPb3Bi, which hosts a nontrivial electronic structure. It exhibits type II superconductivity below 3.01(1) K. Heat capacity and transverse field muon spin rotation relaxation (muSR) measurements demonstrate a fully gapped isotropic s wave state with moderate electron phonon coupling, while zero field muSR confirms the preservation of time reversal symmetry (TRS). Transport measurements reveal low carrier mobility with diffusive normal state transport. Electronic structure calculations show strong dispersion along the quasi 1D direction and relatively flatter bands in the transverse plane, giving rise to pronounced Fermi surface nesting in the kx-ky plane. Consistent with this, the compound undergoes a charge density wave transition at 280(1) K. The flow of Wannier charge centers, together with surface state dispersion, establishes nontrivial band topology. These results identify PtPb3Bi as a new quasi 1D superconductor with nontrivial electronic structure and a promising candidate for topological superconductivity.

cond-mat.supr-con

Hourglass Dirac chains enable intrinsic topological superconductivity in nonsymmorphic silicides

Nonsymmorphic crystalline symmetries provide a robust route to symmetry-protected electronic topology, yet their role in stabilizing intrinsic topological superconductivity remains largely unexplored. Here, we report \ch{TaPtSi} as a new member of the superconducting nonsymmorphic silicide family, characterized via AC transport, magnetization, heat capacity, and muon spin rotation/relaxation ($μ$SR) measurements. Zero field $μ$SR reveals spontaneous internal magnetic fields below $T_{\rm c}$, establishing time reversal symmetry breaking in \ch{TaPtSi}. First principles calculations on \ch{TaPtSi} and its isostructural nonsymmorphic superconducting analogues reveal the presence of symmetry-protected hourglass dispersions. The "necks" of these dispersions form Dirac nodal rings and chains that reside near or intersect the Fermi level. Guided by Ginzburg Landau symmetry analysis, we identify an internally antisymmetric non unitary triplet pairing state as the unique ground state consistent with the experimental phenomenology. Based on Bogoliubov de Gennes calculations, we further demonstrate that this state supports Majorana surface modes, establishing its intrinsically topological nature. These results reveal a systematic route by which nonsymmorphic symmetry drives the interplay between hourglass Dirac chain topology and unconventional triplet pairing, positioning equiatomic silicides as a unified materials platform for intrinsic topological superconductivity.

cond-mat.supr-con

Towards Universal Neural Likelihood Inference

We introduce universal neural likelihood inference (UNLI): enabling a single model to provide data-grounded, conditional likelihood predictions for arbitrary targets given any collection of observed features, across diverse domains and tasks. To achieve UNLI over heterogeneous tabular data, we develop the Arbitrary Set-based Permutation-Invariant Reasoning Engine (ASPIRE) model. Our design addresses critical gaps in existing approaches to merge semantic-understanding capabilities and generalised numerical feature reasoning within a zero-shot capable framework. Trained on over 1,400 real diverse datasets spanning various domains, ASPIRE achieves 15\% higher F1 scores and 85\% lower RMSE than existing tabular foundation models in zero-shot and few-shot settings. Lastly, this work introduces open-world active feature acquisition, where we leverage the UNLI capabilities of ASPIRE to adeptly determine next feature-values to observe to improve inference time prediction accuracies.

cs.LG

DiffVax: Optimization-Free Image Immunization Against Diffusion-Based Editing

Current image immunization defense techniques against diffusion-based editing embed imperceptible noise into target images to disrupt editing models. However, these methods face scalability challenges, as they require time-consuming optimization for each image separately, taking hours for small batches. To address these challenges, we introduce DiffVax, a scalable, lightweight, and optimization-free framework for image immunization, specifically designed to prevent diffusion-based editing. Our approach enables effective generalization to unseen content, reducing computational costs and cutting immunization time from days to milliseconds, achieving a speedup of 250,000x. This is achieved through a loss term that ensures the failure of editing attempts and the imperceptibility of the perturbations. Extensive qualitative and quantitative results demonstrate that our model is scalable, optimization-free, adaptable to various diffusion-based editing tools, robust against counter-attacks, and, for the first time, effectively protects video content from editing. More details are available in https://diffvax.github.io/ .

cs.CV

Observation of Time-Reversal Symmetry Breaking in the Type-I Superconductor YbSb$_2$

The spontaneous breaking of time-reversal symmetry is a hallmark of unconventional superconductivity, typically observed in type-II superconductors. Here, we report evidence of time-reversal symmetry breaking in the type-I superconductor YbSb$_2$. Zero-field $μ$SR measurements reveal spontaneous internal magnetic fields emerging just below the superconducting transition, while transverse-field $μ$SR confirms a fully gapped type-I superconducting state. Our first-principles calculations identify YbSb$_2$ as a ${\mathbb Z}_2$ topological metal hosting a Dirac nodal line near the Fermi level. Symmetry analysis within the Ginzburg Landau framework indicates an internally antisymmetric nonunitary triplet (INT) state as the most probable superconducting ground state. Calculations based on an effective low-energy model further demonstrate that this INT state hosts gapless Majorana surface modes, establishing YbSb$_2$ as a topological superconductor. Our results highlight YbSb$_2$ as a unique material platform where type-I superconductivity coexists with triplet-pairing and nontrivial topology.

cond-mat.supr-con

Socratic Students: Teaching Language Models to Learn by Asking Questions

Large language Models (LLMs) are usually used to answer questions, but many high-stakes applications (e.g., tutoring, clinical support) require the complementary skill of asking questions: detecting missing information, requesting clarifications, and using them to solve tasks. We study this skill in reasoning-heavy domains where progress depends on inquiry rather than factual recall. We define an interactive protocol where a student model engages a stronger teacher under a small turn budget. After each teacher reply, we evaluate the student on the original task with Pass@k. We propose Outcome-Driven Question optimization Strategy (ODQS ), a training framework that learns a questioning policy from downstream task outcomes. At each turn, we sample multiple candidate questions; query the teacher with each, then score the student's resulting performance. Using these scores, we train the student via supervised fine-tuning followed by Direct Preference Optimization (DPO), without any human labels. On GSM8K, HumanEval, and OpenCoder, ODQS produces large gains over interactive baselines, boosting Pass@5 by up to 54.7% (absolute) on math and 22.9% (absolute) on coding, and matching baseline performance in three fewer turns. Thus, question asking can be explicitly trained from task outcomes, improving both accuracy and efficiency in interactive reasoning.

cs.AI

A Causal Lens for Evaluating Faithfulness Metrics

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's true reasoning faithfully. While several faithfulness metrics have been proposed, they are often evaluated in isolation, making principled comparisons between them difficult. We present Causal Diagnosticity, a testbed framework for evaluating faithfulness metrics for natural language explanations. We use the concept of diagnosticity, and employ model-editing methods to generate faithful-unfaithful explanation pairs. Our benchmark includes four tasks: fact-checking, analogy, object counting, and multi-hop reasoning. We evaluate prominent faithfulness metrics, including post-hoc explanation and chain-of-thought methods. Diagnostic performance varies across tasks and models, with Filler Tokens performing best overall. Additionally, continuous metrics are generally more diagnostic than binary ones but can be sensitive to noise and model choice. Our results highlight the need for more robust faithfulness metrics.

cs.CL

Probing the intermediate state of type-I superconductor SnAs using Muon Spin Spectroscopy

Superconductivity with non-trivial band topology provides a novel platform for exploring topological superconductivity and its quantum applications. A detailed microscopic understanding of the superconducting ground state in such materials is crucial. Here, we report the results of a muon spin rotation/relaxation study ($μ$SR) of the topologically non-trivial superconductor SnAs, which exhibits superconductivity below 3.74(1) \si{K}. Zero-field (ZF) $μ$SR data reveal that this system is a time-reversal invariant superconductor, and systematic transverse-field (TF) $μ$SR measurements unveil the type-I nature of the SnAs superconductor. We have established the superconducting phase diagram to understand the intermediate state of type-I superconductors. Moreover, ab \textit{initio} band structure and phonon calculations are performed, which correlate with the experimental characterization.

cond-mat.supr-con

Adversarially Probing Cross-Family Sound Symbolism in 27 Languages

The phenomenon of sound symbolism, the non-arbitrary mapping between word sounds and meanings, has long been demonstrated through anecdotal experiments like Bouba Kiki, but rarely tested at scale. We present the first computational cross-linguistic analysis of sound symbolism in the semantic domain of size. We compile a typologically broad dataset of 810 adjectives (27 languages, 30 words each), each phonemically transcribed and validated with native-speaker audio. Using interpretable classifiers over bag-of-segment features, we find that phonological form predicts size semantics above chance even across unrelated languages, with both vowels and consonants contributing. To probe universality beyond genealogy, we train an adversarial scrubber that suppresses language identity while preserving size signal (also at family granularity). Language prediction averaged across languages and settings falls below chance while size prediction remains significantly above chance, indicating cross-family sound-symbolic bias. We release data, code, and diagnostic tools for future large-scale studies of iconicity.

cs.CL

Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation

LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior. ASR transcripts typically use anonymous labels such as $Speaker\_1$, preventing models from learning stable participant behavior across meetings. We present a reproducible pipeline that converts public Zoom recordings into speaker-attributed transcripts enriched with persona profiles, topics, and pragmatic "action tags" such as $[propose\_motion]$. Using this pipeline, we release three public datasets of government deliberation (Appellate Court hearings, School Board meetings, and Municipal Council sessions) and fine-tune LLM personas on this action-aware data. We evaluate simulations along four dimensions: persona fidelity, persona consistency, institutional fidelity, and behavioral coherence. Action-aware fine-tuning cuts perplexity by 67%, doubles classifier-based persona fidelity, increases vote attempts by up to $3.6\times$, and improves deliberative responsiveness by up to 70%. Human evaluations show that simulated excerpts are often hard to distinguish from real deliberations, indicating a practical foundation for data-grounded civic simulation studies.

cs.CL

From Attention to Disaggregation: Tracing the Evolution of LLM Inference

The evolution of Large Language Models from the Transformer architecture to models with trillions of parameters has shifted the primary bottleneck from model training to real time inference. Deploying these massive models is a complex distributed systems challenge constrained by memory bandwidth, computational throughput, and latency requirements. LLM inference fundamentally requires solving a multi objective optimization problem to minimize latency, maximize throughput, and reduce cost. This paper explores the necessary architectural shift towards disaggregated inference, which applies distributed systems principles such as service decomposition, resource disaggregation, and workload partitioning to overcome the limitations of traditional monolithic GPU clusters. By decoupling the compute intensive prefill phase from the memory intensive decode phase into independently scalable components, this paradigm mitigates resource contention and enables independent optimization of key metrics like Time to First Token and Inter Token Latency.

cs.DC

Nonsymmorphic symmetry protected hourglass Dirac chain topology and conventional superconductivity in ZrIrGe

Ternary transition-metal germanide superconductors with nonsymmorphic symmetries offer promising platforms for symmetry-protected topological phases. In this work, we investigate ZrIrGe, which crystallizes in the nonsymmorphic TiNiSi-type structure. Electrical, magnetic, and specific heat measurements confirm bulk type-II superconductivity with a full gap and a transition temperature of 2.84(7) K, consistent with weak-coupling BCS behavior. First-principles calculations reveal hourglass-shaped bulk band dispersions and a Dirac chain composed of symmetry-protected fourfold-degenerate Dirac points, leading to drumhead-like surface states near the Fermi level. Additionally, ZrIrGe exhibits a nontrivial $\mathbb{Z}_2$ topological character, resulting in helical surface states that cross the Fermi level, making it a strong candidate for proximity-induced topological superconductivity. The coexistence of conventional superconductivity and topological band features establishes ZrIrGe as a rare stoichiometric system for exploring intrinsic topological superconductivity.

cond-mat.supr-con

Algorithmic Improvements to List Decoding of Folded Reed-Solomon Codes

Folded Reed-Solomon (FRS) codes are a well-studied family of codes, known for achieving list decoding capacity. In this work, we give improved deterministic and randomized algorithms for list decoding FRS codes of rate $R$ up to radius $1-R-\varepsilon$. We present a deterministic decoder that runs in near-linear time $\widetilde{O}_{\varepsilon}(n)$, improving upon the best-known runtime $n^{Ω(1/\varepsilon)}$ for decoding FRS codes. Prior to our work, no capacity achieving code was known whose deterministic decoding could be done in time $\widetilde{O}_{\varepsilon}(n)$. We also present a randomized decoder that runs in fully polynomial time $\mathrm{poly}(1/\varepsilon) \cdot \widetilde{O}(n)$, improving the best-known runtime $\mathrm{exp}(1/\varepsilon)\cdot \widetilde{O}(n)$ for decoding FRS codes. Again, prior to our work, no capacity achieving code was known whose decoding time depended polynomially on $1/\varepsilon$. Our results are based on improved pruning procedures for finding the list of codewords inside a constant-dimensional affine subspace.

cs.IT