SearcharxivSearch

arXiv subjects

Yi-Cheng Wang

Publications and source records attributed to Yi-Cheng Wang.

At least 19 recordsLinked to original sources

Multimodal Graph RAG for Long-range Visually Rich Document Understanding

Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue by the limited context window. Though recent multimodal retrieval-augmented generation (MMRAG) can address this challenge by retrieving relevant pages. It still struggles with the visual question answering (VQA) requiring holistic comprehension of a document. To cope with this, knowledge graph (KG) that summarizes global knowledge of a document can provide an effective solution. However, most existing LLM-based KG construction methods handle only the language modality, leaving the automatic creation of multimodal KGs (MMKGs) for visually rich documents largely unexplored. In this paper, we introduce a multimodal graph-based RAG approach to tackle this problem. Existing LLM-based KG methods evaluate the QA performance relying on indirect evidence such as comprehensiveness, diversity, empowerment, and so on. The lack of annotated datasets for comprehensive document-level VQA poses a significant challenge to effective model evaluation. To overcome this limitation, we also introduce a new benchmark, DLVQA (document-level VQA), which provides reference summaries and corresponding supporting facts for global document-level questions. Experimental results show that our approach outperforms existing MMRAG or KG-based approaches on multi-hop QA/VQA benchmarks and DLVQA.

cs.IR

Entanglement transitions in translation-invariant tensor networks

We study the complexity of approximately contracting translation-invariant tensor networks. The computational cost of row-by-row tensor network contraction, which defines a discrete time evolution governed by a fixed transfer matrix, is associated with the entanglement of the state of a row. By analyzing a family of tensor networks whose transfer matrices interpolate between chaotic Floquet and strongly non-unitary limits, we uncover a transition between volume- and area-law entanglement in states evolved under the transfer matrix. We show that deep in the volume-law phase the spectrum of the transfer matrix in the complex plane consists of a dense ring with a sharp outer edge, reminiscent of behavior identified for non-unitary random matrices. At late times an evolving row state therefore has significant contributions from many eigenvectors with nearly degenerate eigenvalue magnitudes. In the area-law phase, there is instead a distinct leading eigenvalue. Our results establish connections between contraction complexity, spectral properties of the transfer matrix, and purification under non-unitary dynamics.

quant-ph

MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation

Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unseen documents. Nonetheless, they struggle with high-level conceptual understanding and holistic comprehension due to limited context windows, which constrain their ability to perform deep reasoning over long-form, domain-specific content such as full-length books. To solve this problem, knowledge graphs (KGs) have been leveraged to provide entity-centric structure and hierarchical summaries, offering more structured support for reasoning. However, existing KG-based RAG solutions remain restricted to text-only inputs and fail to leverage the complementary insights provided by other modalities such as vision. On the other hand, reasoning from visual documents requires textual, visual, and spatial cues into structured, hierarchical concepts. To address this issue, we introduce a multimodal knowledge graph-based RAG that enables cross-modal reasoning for better content understanding. Our method incorporates visual cues into the construction of knowledge graphs, the retrieval phase, and the answer generation process. Experimental results across both global and fine-grained question answering tasks show that our approach consistently outperforms existing RAG-based approaches on both textual and multimodal corpora.

cs.AI

Document-Level Numerical Reasoning across Single and Multiple Tables in Financial Reports

Despite the strong language understanding abilities of large language models (LLMs), they still struggle with reliable question answering (QA) over long, structured documents, particularly for numerical reasoning. Financial annual reports exemplify this difficulty: financial statement analysis often hinges on accurate arithmetic, and analysts derive key indicators by integrating evidence scattered across multiple tables and narrative text. However, existing benchmarks focus largely on single-table settings, leaving cross-table document-level numerical reasoning underexplored. To address this gap, we introduce FinLongDocQA, a dataset for both single-table and cross-table financial numerical reasoning in long-context reports. Evaluating both closed-source and open-source LLMs on FinLongDocQA reveals two bottlenecks: (1) annual reports often exceed 129k tokens, exacerbating the context rot problem for locating relevant tables; and (2) even when relevant evidence is located, LLMs remain prone to errors in multi-step numerical reasoning. We propose FinLongDocAgent, a Multi-Agent Multi-Round Retrieval-Augmented Generation (RAG) approach that iteratively retrieves evidence, performs intermediate calculations, and verifies results across rounds. Experiments highlight the importance of iterative retrieval and verification for reliable numerical QA in long financial documents.

cs.CL

Stability of quantum chaos against weak non-unitarity

We study the quantum dynamics generated by the repeated action of a non-unitary evolution operator on a system of qubits. Breaking unitarity can lead to the purification of mixed initial states, which corresponds to the loss of sensitivity to initial conditions, and hence the absence of a key signature of dynamical chaos. However, the scrambling of quantum information can delay purification to times that are exponential in system size. Here we study purification in systems whose evolution operators are fixed in time, where all aspects of the dynamics are in principle encoded in spectral properties of the evolution operator for a single time step. The operators that we study consist of global Haar random unitary operators and non-unitary single-qubit operations. We show that exponentially slow purification arises from a distribution of eigenvalues in the complex plane that forms a ring with sharp edges at large radii, with the eigenvalue density exponentially large near these edges. We argue that the sharp edges of the eigenvalue distribution arise from level attraction along the radial direction in the complex plane. By calculating the spectral form factor we also show that there is level repulsion around the azimuthal direction, even close to the outer edge of the ring of eigenvalues. Our results connect this spectral signature of quantum chaos to the sensitivity of the system to its initial conditions.

quant-ph

Vacuum Rabi Splitting and Quantum Fisher Information of a Non-Hermitian Qubit in a Single-Mode Cavity

A natural extension of the non-Hermitian qubit is to place it in a single-mode cavity. This setup corresponds to the quantum Rabi model (QRM) with a purely imaginary bias on the qubit, exhibiting parity-time ($\mathcal{P}\mathcal{T}$) symmetry. In this work, we first solve the $\mathcal{P} \mathcal{T}$-symmetric QRM using the Bogoliubov operator approach. We derive the transcendental function responsible for the exact solution, which can also be used to precisely identify exceptional points. The adiabatic approximation previously used can be easily formulated within this approach by considering transitions between the same manifolds in the space of Bogoliubov operators. By further considering transitions between the nearest-neighboring manifolds, we can analytically obtain more accurate eigensolutions. Moreover, these simple corrections can capture the main features of the dynamics, where the adiabatic approximation fails. Furthermore, the rich characteristics of the vacuum Rabi splitting in the emission spectrum are predicted. The width of the peaks increases with the coupling strength and the imaginary biases, reflecting the nature of open quantum systems. Additionally, we identify a {quantum-criticality-enhanced} effect by calculating the quantum Fisher information. Near the exceptional points, the quantum Fisher information in the $\mathcal{P} \mathcal{T}$-symmetric QRM is significantly higher than that of the non-Hermitian qubit component. This may open a new avenue for enhancing quantum sensitivity in non-Hermitian systems by incorporating coupling with an additional degree of freedom, enabling more precise parameter estimation.

quant-ph

Multi-task Pretraining for Enhancing Interpretable L2 Pronunciation Assessment

Automatic pronunciation assessment (APA) analyzes second-language (L2) learners' speech by providing fine-grained pronunciation feedback at various linguistic levels. Most existing efforts on APA typically adopt segmental-level features as inputs and predict pronunciation scores at different granularities via hierarchical (or parallel) pronunciation modeling. This, however, inevitably causes assessments across linguistic levels (e.g., phone, word, and utterance) to rely solely on phoneme-level pronunciation features, nearly sidelining supra-segmental pronunciation cues. To address this limitation, we introduce multi-task pretraining (MTP) for APA, a simple yet effective strategy that attempts to capture long-term temporal pronunciation cues while strengthening the intrinsic structures within an utterance via the objective of reconstructing input features. Specifically, for a phoneme-level encoder of an APA model, the proposed MTP strategy randomly masks segmental-level pronunciation features and reconstructs the masked ones based on their surrounding pronunciation context. Furthermore, current APA systems lack integration with automated speaking assessment (ASA), limiting holistic proficiency evaluation. Drawing on empirical studies and prior knowledge in ASA, our framework bridges this gap by incorporating handcrafted features (HCFs), such as fluency (speech rate, silence duration) and stress (pitch accent strength), derived from human-designed formulas via regressors to generate interpretable proficiency scores. Experiments on speechocean762 show improved pronunciation scoring and ASA proficiency correlation, enabling targeted training and comprehensive proficiency assessment.

cs.CL

Towards robust variational quantum simulation of Lindblad dynamics via stochastic Magnus expansion

In this paper, we introduce a novel and general framework for the variational quantum simulation of Lindblad equations. Building on the close relationship between the unraveled Lindblad dynamics, stochastic Magnus integrators, and variational quantum simulation, we propose a high-order scheme for solving the quantum state diffusion equation using exponential integrators. This formulation facilitates the simulation of wavefunction trajectories within the established framework of variational quantum algorithms for time evolution. Our algorithm significantly enhances robustness in two key aspects: the stability of the simulation with large time steps, and the reduction in the number of quantum trajectories required to accurately simulate the Lindblad dynamics in terms of the ensemble average. We demonstrate the effectiveness of our algorithm through numerical examples in both classical and quantum implementations, including the transverse-field Ising model (TFIM) with damping, the Fenna-Matthews-Olson (FMO) complex, and the radical pair model (RPM). The simulation accuracy can be systematically improved, and the algorithm remains reliable even in highly oscillatory regimes. These methods are expected to be applicable to a broader class of open quantum systems beyond the specific models considered in this study.

quant-ph

$\mathcal{PT}$-symmetric two-photon quantum Rabi models

We investigate two non-Hermitian two-photon quantum Rabi models (tpQRM) that exhibit $\mathcal{PT}$ symmetry: the biased tpQRM (btpQRM), in which the qubit bias is purely imaginary, and the dissipative tpQRM (dtpQRM), where the two-photon coupling is made imaginary to introduce dissipation. For both models, we derive exact solutions by employing Bogoliubov transformations. In the btpQRM, we identify spectral collapse at a critical coupling strength, with accompanying $\mathcal{PT}$ symmetry breaking that correlates with exceptional points (EPs) arising from coalescing eigenstates. We establish a direct correspondence between $\mathcal{PT}$-broken regions and the doubly degenerate points of the Hermitian tpQRM, and analyze the effects of qubit bias via an adiabatic approximation. In the dtpQRM, although no spectral collapse occurs, both EPs and Juddian-type degeneracies are present, with well-separated behaviors distinguished by parity conservation. Through biorthogonal fidelity susceptibility and c-product, we successfully identify and classify the nature of these two types of level crossings. Finally, we compare the dynamical evolution of both models, revealing fundamentally different pathways to steady states governed by their respective non-Hermitian spectral structures. Our results provide exact characterizations of $\mathcal{PT}$-symmetric non-Hermitian tpQRMs and may offer theoretical insights for future experimental realizations.

quant-ph

$\mathcal{PT}$-symmetric quantum Rabi model: Solutions and exceptional points

The $\mathcal{PT}$-symmetric non-Hermitian quantum Rabi model (QRM) with imaginary coupling is solved using the Bogoliubov operators approach. A transcendental function responsible for the exact solutions is derived, with its zeros yielding the regular spectrum. We find two types of intersections: One is the exceptional point (EP), which is widely studied in the non-Hermitian system; another one is due to doubly degenerate states caused by the conserved QRM parity, which is well-known in the Hermitian QRM. These intersections are identified through this transcendental function. EPs emerge between pairs of adjacent excited energy levels, shifting toward lower coupling strengths as energy levels increase. The fidelity susceptibility diverges to negative infinity at the EPs, consistent with recent findings in non-Hermitian systems, while it diverges to positive infinity at the doubly degenerate points. The EPs are further confirmed by the vanishing c-product in the biorthogonal basis. All eigenstates are characterized by conserved energy and QRM parity. We conclude that the non-Hermitian QRM is integrable, analogous to its Hermitian counterpart.

quant-ph

Dissipative quantum phase transitions in electrically driven lasers

Embedding quantum dot circuits into microwave cavities has emerged as a novel platform for controlling photon emission statistics by electrical means. With such a circuit version of the Rabi model, we reveal previously undefined quantum phase transitions in electrically driven lasing regimes, which do not require deep strong light-matter couplings. For one-photon interaction, the scaling analysis indicates that the system undergoes a continuous phase transition from thermal to coherent photon emissions. Going beyond this, a discontinuous quantum phase transition from superbunched to coherent states in two-photon processes, accompanied by the bistability within a mean-field theory, is predicted. Both the order of phase transitions and the critical electron-photon coupling can be easily controlled by an electric field, while the tunneling current can be used as a fingerprint of such transitions. Our prediction, along with its extension to multiphoton processes, represents a key step towards accessing lasing phase transitions.

cond-mat.mes-hall

Automated Speaking Assessment of Conversation Tests with Novel Graph-based Modeling on Spoken Response Coherence

Automated speaking assessment in conversation tests (ASAC) aims to evaluate the overall speaking proficiency of an L2 (second-language) speaker in a setting where an interlocutor interacts with one or more candidates. Although prior ASAC approaches have shown promising performance on their respective datasets, there is still a dearth of research specifically focused on incorporating the coherence of the logical flow within a conversation into the grading model. To address this critical challenge, we propose a hierarchical graph model that aptly incorporates both broad inter-response interactions (e.g., discourse relations) and nuanced semantic information (e.g., semantic words and speaker intents), which is subsequently fused with contextual information for the final prediction. Extensive experimental results on the NICT-JLE benchmark dataset suggest that our proposed modeling approach can yield considerable improvements in prediction accuracy with respect to various assessment metrics, as compared to some strong baselines. This also sheds light on the importance of investigating coherence-related facets of spoken responses in ASAC.

cs.CL

Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection

Code-switching-where multilingual speakers alternately switch between languages during conversations-still poses significant challenges to end-to-end (E2E) automatic speech recognition (ASR) systems due to phenomena of both acoustic and semantic confusion. This issue arises because ASR systems struggle to handle the rapid alternation of languages effectively, which often leads to significant performance degradation. Our main contributions are at least threefold: First, we incorporate language identification (LID) information into several intermediate layers of the encoder, aiming to enrich output embeddings with more detailed language information. Secondly, through the novel application of language boundary alignment loss, the subsequent ASR modules are enabled to more effectively utilize the knowledge of internal language posteriors. Third, we explore the feasibility of using language posteriors to facilitate deep interaction between shared encoder and language-specific encoders. Through comprehensive experiments on the SEAME corpus, we have verified that our proposed method outperforms the prior-art method, disentangle based mixture-of-experts (D-MoE), further enhancing the acuity of the encoder to languages.

eess.AS

An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition

End-to-end (E2E) automatic speech recognition (ASR) models have become standard practice for various commercial applications. However, in real-world scenarios, the long-tailed nature of word distribution often leads E2E ASR models to perform well on common words but fall short in recognizing uncommon ones. Recently, the notion of a contextual adapter (CA) was proposed to infuse external knowledge represented by a context word list into E2E ASR models. Although CA can improve recognition performance on rare words, two crucial data imbalance problems remain. First, when using low-frequency words as context words during training, since these words rarely occur in the utterance, CA becomes prone to overfit on attending to the token due to higher-frequency words not being present in the context list. Second, the long-tailed distribution within the context list itself still causes the model to perform poorly on low-frequency context words. In light of this, we explore in-depth the impact of altering the context list to have words with different frequency distributions on model performance, and meanwhile extend CA with a simple yet effective context-balanced learning objective. A series of experiments conducted on the AISHELL-1 benchmark dataset suggests that using all vocabulary words from the training corpus as the context list and pairing them with our balanced objective yields the best performance, demonstrating a significant reduction in character error rate (CER) by up to 1.21% and a more pronounced 9.44% reduction in the error rate of zero-shot words.

cs.CL

Light scattering properties beyond weak-field excitation in atomic ensembles

In the study of optical properties of large atomic system, a weak laser driving is often assumed to simplify the system dynamics by linearly coupled equations. Here, we investigate the light scattering properties of atomic ensembles beyond weak-field excitation through the cumulant expansion method. By progressively incorporating higher-order correlations into the steady-state equations, an enhanced accuracy can be achieved in comparison to the exact solutions from solving a full density matrix. Our analysis reveals that, in the regime of weak dipole-dipole interaction (DDI), the first-order expansion yields satisfactory predictions for optical depth, while denser atomic configurations necessitate consideration of higher-order correlations. As the intensity of incident light increases, atom saturation effects become noticeable, giving rise to significant changes in light transparency, energy shift, and decay rate. This saturation phenomenon extends to subradiant atom arrays even under weak driving conditions, leading to substantial deviations from the linear model. Our findings demonstrate the mean-field models as good extensions to linear models as it balances both accuracy and computational complexity. However, the crucial role of higher-order cumulants in large and dense atom systems remains unclear, since it is challenging theoretically owing to the exponentially increasing Hilbert space in such light-matter interacting systems.

quant-ph

ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization

Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, where the models are trained to estimate target values without explicitly accounting for phoneme-awareness in the feature space. In this paper, we propose a contrastive phonemic ordinal regularizer (ConPCO) tailored for regression-based APA models to generate more phoneme-discriminative features while considering the ordinal relationships among the regression targets. The proposed ConPCO first aligns the phoneme representations of an APA model and textual embeddings of phonetic transcriptions via contrastive learning. Afterward, the phoneme characteristics are retained by regulating the distances between inter- and intra-phoneme categories in the feature space while allowing for the ordinal relationships among the output targets. We further design and develop a hierarchical APA model to evaluate the effectiveness of our method. Extensive experiments conducted on the speechocean762 benchmark dataset suggest the feasibility and efficacy of our approach in relation to some cutting-edge baselines.

eess.AS

DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition

End-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently been proposed, which normally build on phonetic-level edit distance algorithms and have shown impressive NEC performance. However, as the named entity (NE) list grows, the problems of phonetic confusion in the NE list are exacerbated; for example, homophone ambiguities increase substantially. In view of this, we proposed a novel Description Augmented Named entity CorrEctoR (dubbed DANCER), which leverages entity descriptions to provide additional information to facilitate mitigation of phonetic confusion for NEC on ASR transcription. To this end, an efficient entity description augmented masked language model (EDA-MLM) comprised of a dense retrieval model is introduced, enabling MLM to adapt swiftly to domain-specific entities for the NEC task. A series of experiments conducted on the AISHELL-1 and Homophone datasets confirm the effectiveness of our modeling approach. DANCER outperforms a strong baseline, the phonetic edit-distance-based NEC model (PED-NEC), by a character error rate (CER) reduction of about 7% relatively on AISHELL-1 for named entities. More notably, when tested on Homophone that contain named entities of high phonetic confusion, DANCER offers a more pronounced CER reduction of 46% relatively over PED-NEC for named entities.

cs.CL

An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement

With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the codeswitching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the variations between languages often lead to degradation of ASR performance. In this paper, we focus exclusively on improving the acoustic encoder of E2E ASR to tackle the challenge caused by the codeswitching phenomenon. Our main contributions are threefold: First, we introduce a novel disentanglement loss to enable the lower-layer of the encoder to capture inter-lingual acoustic information while mitigating linguistic confusion at the higher-layer of the encoder. Second, through comprehensive experiments, we verify that our proposed method outperforms the prior-art methods using pretrained dual-encoders, meanwhile having access only to the codeswitching corpus and consuming half of the parameterization. Third, the apparent differentiation of the encoders' output features also corroborates the complementarity between the disentanglement loss and the mixture-of-experts (MoE) architecture.

cs.CL