SearcharxivSearch

arXiv subjects

Navya Gupta

Publications and source records attributed to Navya Gupta.

11 recordsLinked to original sources

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

Automatic speech recognition (ASR) performance degrades sharply on dysarthric speech, yet how disordered articulation reshapes a model's internal representations is underexplored. We conduct a layer-wise probing analysis of a transformer ASR encoder on Mandarin dysarthric speech under three transcript-matched conditions: original dysarthric speech, speaker conditioned zero-shot TTS resynthesis, and unconditioned TTS. Probing reveals a task- and condition-dependent representation hierarchy: phoneme boundary information remains weak across all layers for dysarthric speech; phoneme identity is recoverable in deep layers for synthetic speech, but remains poor for dysarthric speech; and recognition difficulty is concentrated in the deepest layers. Furthermore, lexical tone is a persistent error source across all conditions. Guided by these insights, layer-selective LoRA shows that mid-layer adaptation (layer 7 or layers 5-8) recovers near-full encoder performance on dysarthric speech within 6.67% and 2.89% relative margins while training only 0.16% and 0.65% of adapter parameters. Conversely, upper-layer adaptation benefits synthetic speech more than dysarthric speech. These findings link representation analysis to parameter-efficient fine-tuning and motivate layer-aware adaptation for low-resource Mandarin dysarthric ASR.

cs.CL

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational pathways through which they arise and propagate. We introduce an Operation-centric mechanistic framework that decomposes VLM failures by both the reasoning operation where they originate and the internal computational pathway through which they propagate. Our analysis reveals four dominant failure modes: grounding failure, reasoning failure, attribute extraction failure, and language-prior dominance, each characterized by a distinct relationship between visual grounding strength and answer correctness. Through three complementary causal interventions applied across all transformer layers, we find that object-selection failures are associated primarily with feedforward computation, multi-step relational failures with late-layer direct attention, and attribute-extraction failures with answer-position feedforward computation. Validation on VSR further shows that single-step spatial failures are concentrated at object-position encoding, distinguishing them from multi-step relational composition. These findings reveal distinct computational bottlenecks across operation types and provide a principled basis for targeted diagnosis of VLM failures in multimedia reasoning.

cs.CV

IRIS: Interleaved Reinforcement with Incremental Staged Curriculum for Cross-Lingual Mathematical Reasoning

Curriculum learning helps language models tackle complex reasoning by gradually increasing task difficulty. However, it often fails to generate consistent step-by-step reasoning, especially in multilingual and low-resource settings where cross-lingual transfer from English to Indian languages remains limited. We propose IRIS: Interleaved Reinforcement with Incremental Staged Curriculum, a two-axis framework that combines Supervised Fine-Tuning on progressively harder problems (vertical axis) with Reverse Curriculum Reinforcement Learning to reduce reliance on step-by-step guidance (horizontal axis). We design a composite reward combining correctness, step-wise alignment, continuity, and numeric incentives, optimized via Group Relative Policy Optimization (GRPO). We release CL-Math, a dataset of 29k problems with step-level annotations in English, Hindi, and Marathi. Across standard benchmarks and curated multilingual test sets, IRIS consistently improves performance, with strong results on math reasoning tasks and substantial gains in low-resource and bilingual settings, alongside modest improvements in high-resource languages.

cs.CL

String-breaking statics and dynamics in a (1+1)D SU(2) lattice gauge theory

String breaking is at the core of hadronization models of relevance to particle colliders. Yet, studies of string-breaking dynamics rooted in quantum chromodynamics remain fundamentally challenging. Tensor networks enable sign-problem-free studies of static and dynamical properties of lattice gauge theories. In this work, we develop and apply a tensor-network toolkit based on the loop-string-hadron formulation of an SU(2) lattice gauge theory in 1+1 dimensions with dynamical fermions. We apply this toolkit to study static and dynamical aspects of strings and their breaking in this theory. The simple, gauge-invariant, and local structure of the loop-string-hadron states and constraints removes the need to impose non-Abelian constraints in the algorithm, and allows for a systematic computation of observables at increasingly large bosonic cutoffs, and toward the infinite-volume and continuum limits. Our study of static strings yields a determination of the string tension in the continuum and thermodynamic limits. Our study of dynamical string breaking, performed at a fixed lattice spacing and system size, illuminates underlying processes at play during the quench dynamics of a string. The loop, string, and hadron description offers a systematic and intuitive way to diagnose these processes, including string expansion and contraction, endpoint splitting and particle shower, chain scattering events, and inelastic processes resulting from string dissociation and recombination, and particle production. We relate these processes to several features of the dynamics, such as energy transport, entanglement-entropy production, and correlation spreading. This work opens the way to future tensor-network studies of string breaking and particle production in increasingly complex lattice gauge theories.

hep-lat

A Euclidean Monte-Carlo-informed route to ground-state preparation for quantum simulation of scalar field theory

Quantum simulators hold great promise for studying real-time (Minkowski) dynamics of quantum field theories. Nonetheless, preparing non-trivial initial states remains a major obstacle. Euclidean-time Monte-Carlo methods yield ground-state spectra and static correlation functions that can, in principle, guide state preparation. In this work, we exploit this classical information to bridge Euclidean and Minkowski descriptions for a (1+1)-dimensional interacting scalar field theory. We propose variational ansatz families which achieve comparable ground-state energies, yet exhibit distinct correlations and local non-Gaussianity. By optimizing selected wavefunction moments with Monte-Carlo data, we obtain ansatzes that can be efficiently translated into quantum circuits. Our algorithmic cost analysis shows these circuits' gate complexity scales polynomially in system size. Our work paves the way for systematically leveraging classically-computed information to prepare initial states in quantum field theories of interest in nature.

hep-lat

Euclidean-Monte-Carlo-informed ground-state preparation for quantum simulation of scalar field theory

Quantum simulators offer great potential for investigating dynamical properties of quantum field theories. However, preparing accurate non-trivial initial states for these simulations is challenging. Classical Euclidean-time Monte-Carlo methods provide a wealth of information about states of interest to quantum simulations. Thus, it is desirable to facilitate state preparation on quantum simulators using this information. To this end, we present a fully classical pipeline for generating efficient quantum circuits for preparing the ground state of an interacting scalar field theory in 1+1 dimensions. The first element of this pipeline is a variational ansatz family based on the stellar hierarchy for bosonic quantum systems. The second element of this pipeline is the classical moment-optimization procedure that augments the standard variational energy minimization by penalizing deviations in selected sets of ground-state correlation functions (i.e., moments). The values of ground-state moments are sourced from classical Euclidean methods. The resulting states yield comparable ground-state energy estimates but exhibit distinct correlations and local non-Gaussianity. The third element of this pipeline is translating the moment-optimized ansatz into an efficient quantum circuit with an asymptotic cost that is polynomial in system size. This work opens the way to systematically applying classically obtained knowledge of states to prepare accurate initial states in quantum field theories of interest in nature.

quant-ph

Tensor-network toolbox for probing dynamics of non-Abelian gauge theories

Tensor-network methods enable probing dynamics of strongly interacting quantum many-body systems, including gauge theories, via Hamiltonian simulation, hence bypassing sign problems. They also have the potential to inform efficient quantum-simulation algorithms of the same theories. We develop and benchmark a matrix-product-state ansatz for the SU(2) lattice gauge theory using the loop-string-hadron formulation. This formulation has been demonstrated to be advantageous in Hamiltonian simulation of non-Abelian gauge theories. It is applicable to both SU(2) and SU(3) gauge groups, to periodic and open boundary conditions, and to 1+1 and higher dimensions. In this work, we report on progress in computing static and dynamical observables in a SU(2) gauge theory in (1+1)D, pushing the boundary of existing studies.

hep-lat

Steps are all you need: Rethinking STEM Education with Prompt Engineering

Few shot and Chain-of-Thought prompting have shown promise when applied to Physics Question Answering Tasks, but are limited by the lack of mathematical ability inherent to LLMs, and are prone to hallucination. By utilizing a Mixture of Experts (MoE) Model, along with analogical prompting, we are able to show improved model performance when compared to the baseline on standard LLMs. We also survey the limits of these prompting techniques and the effects they have on model performance. Additionally, we propose Analogical CoT prompting, a prompting technique designed to allow smaller, open source models to leverage Analogical prompting, something they have struggled with, possibly due to a lack of specialist training data.

cs.CL

Knowledge Graphs are all you need: Leveraging KGs in Physics Question Answering

This study explores the effectiveness of using knowledge graphs generated by large language models to decompose high school-level physics questions into sub-questions. We introduce a pipeline aimed at enhancing model response quality for Question Answering tasks. By employing LLMs to construct knowledge graphs that capture the internal logic of the questions, these graphs then guide the generation of subquestions. We hypothesize that this method yields sub-questions that are more logically consistent with the original questions compared to traditional decomposition techniques. Our results show that sub-questions derived from knowledge graphs exhibit significantly improved fidelity to the original question's logic. This approach not only enhances the learning experience by providing clearer and more contextually appropriate sub-questions but also highlights the potential of LLMs to transform educational methodologies. The findings indicate a promising direction for applying AI to improve the quality and effectiveness of educational content.

cs.CL

Einstein-Cartan-Dirac equations in the Newman-Penrose formalism

We formulate the Einstein-Cartan-Dirac equations in the Newman-Penrose (NP) formalism, thereby presenting a more accurate and explicit analysis of previous such studies. The equations show in a transparent way how the Einstein-Dirac equations are modified by the inclusion of torsion. In particular, the Hehl-Datta equation is presented in NP notation. We then describe a few solutions of the Hehl-Datta equation on Minkowski space-time, and in particular report a solitonic solution which removes the unphysical behavioiur of the corresponding Dirac solution. The present work serves as a prelude to similar studies for non-degenerate Poincare gauge gravity.

gr-qc

Demonstration of the No-Hiding Theorem on the 5 Qubit IBM Quantum Computer in a Category Theoretic Framework

Quantum no-Hiding theorem, first proposed by Braunstein and Pati [Phys. Rev. Lett. 98, 080502 (2007)], was verified experimentally by Samal et al. [Phys. Rev. Lett. 186, 080401 (2011)] using NMR quantum processor. Till then, this fundamental test has not been explored in any of the experimental architecture. Here, we demonstrate the above no-hiding theorem using the IBM 5Q quantum processor. Categorical algebra developed by Coecke and Duncan [New J. Phys. 13, 043016 (2011)] has been used for better visualization of the no-hiding theorem by analyzing the quantum circuit using the ZX calculus. The experimental results confirm the recovery of missing information by the application of local unitary operations on the ancillary qubits.

quant-ph