SearcharxivSearch

arXiv subjects

Xinxin Li

Publications and source records attributed to Xinxin Li.

At least 19 recordsLinked to original sources

Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations

Automatic speech recognition (ASR) correction has traditionally focused on isolated utterances or short local contexts. However, as text and speech become increasingly interleaved in long interactions, ASR correction requires conversation-level contextual evidence. Existing ASR correction methods often rely on the current hypothesis or concatenate raw dialogue history. In such contexts, sparse correction evidence can be difficult to locate amid redundancy and noise. Addressing these challenges, we propose an ontology memory-augmented ASR correction framework for long text-speech interleaved conversations. The framework organizes preceding interaction history into a dynamically updatable ontology memory, where entities, terminology, surface variants, potential ASR confusions, and semantic relations are stored as retrievable nodes for context-grounded correction. To evaluate this setting, we construct RAMC-Corr, a dataset derived from MAGIC-RAMC for long-range ASR correction with grounded context. Experiments on RAMC-Corr show that our method improves over direct correction in 9 out of 10 paired backbone-setting combinations and encourages more selective and evidence-grounded corrections for context-dependent ASR errors.

cs.CL

EditSR: Enhancing Neural Symbolic Regression via Edit-based Rectification

Neural symbolic regression models improve inference efficiency by shifting structural search to pretraining, but their one-pass autoregressive decoding is prone to error accumulation, which may lead to generating structurally incorrect expressions, especially in complex expression generation scenarios. Existing rectification strategies can alleviate this issue, but they often depend on restarting global search, thereby weakening the efficiency advantage of neural models, and remain susceptible to error accumulation. In this paper, we propose EditSR, a two-layer framework that combines a neural symbolic regression model in the first layer with an edit-based Rectifier in the second layer to achieve efficient prediction and post-hoc rectification. Instead of restarting the global search, we maintain rectification efficiency by pretraining the Rectifier. Specifically, we formulate the rectification process as a step-by-step state-transition chain starting from an incorrect expression, and develop a state-transition algorithm to construct supervised rectification chains for training the Rectifier. To ensure syntactic validity throughout rectification, each edit action is restricted to a syntactically valid space so that every edited expression remains parseable. In addition, because each edit decision is conditioned on the current state rather than the history, the Rectifier allows errors made in earlier steps to be rectified by subsequent edits, thereby reducing the risk of error accumulation. Extensive experiments and ablation studies show that EditSR substantially improves symbolic structure recovery with limited extra cost, with more pronounced gains on complex expressions, where one-pass autoregressive decoding is more susceptible to error accumulation.

cs.AI

Toward Realistic Wi-Fi Fault Diagnosis: A Multi-Modal Benchmark

Intelligent network operation and maintenance systems in modern networks continuously generate large volumes of multi-modal operational data. However, Wi-Fi fault diagnosis under heterogeneous operational environments remains insufficiently understood. We build a real-world Wi-Fi testbed deployed in campus working environments with an automated fault injection system, and collect a multi-modal Wi-Fi fault dataset containing over 10,000 fault samples across diverse wireless scenarios. To the best of our knowledge, this is among the first publicly available datasets jointly capturing heterogeneous cross-layer operational observations for Wi-Fi fault diagnosis. Based on this dataset, we establish a unified benchmark spanning multiple diagnosis tasks, operational modalities, and representative diagnosis paradigms. Experimental results indicate that effectively leveraging heterogeneous operational data remains challenging for existing diagnosis approaches. We further evaluate emerging LLM-based approaches and develop a reasoningoriented evaluation framework to assess the consistency between generated diagnostic analyses and actual network conditions. Our findings suggest several important considerations for future multi-modal Wi-Fi diagnosis.

cs.NI

Generalized Composed Alternating Relaxed Projection Algorithm for Two-Set Feasibility Problem

We study the two-set feasibility problem of finding a point in the intersection $X\cap Y$ of closed convex sets in a Hilbert space. We propose a generalized composed alternating relaxed projection algorithm (gCARPA) that blends Douglas-Rachford-type and projection-reflection-type dynamics via an outer averaging step $\mu$ and an internal relaxation $(\gamma,\theta,\eta)$. The algorithm contains several classical projection methods as special cases. We also introduce its non-stationary variant, in which $(\gamma_k,\theta_k,\eta_k)$ vary over iterations, and establish its convergence. For the subspace feasibility model, we derive an explicit spectral characterization via principal-angle block decompositions, yielding computable subdominant-eigenvalue factors and a minimax parameter-selection recipe in a symmetric regime that targets critical damping on principal-angle planes. Numerical experiments illustrate that the generalized relaxation and its non-stationary tuning can improve or match baseline methods in problem-dependent regimes.

math.OC

Weak-PDE-Net: Discovering Open-Form PDEs via Differentiable Symbolic Networks and Weak Formulation

Discovering governing Partial Differential Equations (PDEs) from sparse and noisy data is a challenging issue in data-driven scientific computing. Conventional sparse regression methods often suffer from two major limitations: (i) the instability of numerical differentiation under sparse and noisy data, and (ii) the restricted flexibility of a pre-defined candidate library. We propose Weak-PDE-Net, an end-to-end differentiable framework that can robustly identify open-form PDEs. Weak-PDE-Net consists of two interconnected modules: a forward response learner and a weak-form PDE generator. The learner embeds learnable Gaussian kernels within a lightweight MLP, serving as a surrogate model that adaptively captures system dynamics from sparse observations. Meanwhile, the generator integrates a symbolic network with an integral module to construct weak-form PDEs, avoiding explicit numerical differentiation and improving robustness to noise. To relax the constraints of the pre-defined library, we leverage Differentiable Neural Architecture Search strategy during training to explore the functional space, which enables the efficient discovery of open-form PDEs. The capability of Weak-PDE-Net in multivariable systems discovery is further enhanced by incorporating Galilean Invariance constraints and symmetry equivariance hypotheses to ensure physical consistency. Experiments on several challenging PDE benchmarks demonstrate that Weak-PDE-Net accurately recovers governing equations, even under highly sparse and noisy observations.

cs.LG

Topological Metamaterial for Magnetic Resonance Imaging

Magnetic Resonance Imaging (MRI) is crucial in global healthcare, but the traditional receive coils, as a core component of MRI, SNR enhancement is limited due to the optimization of channel number and magnetic field strength faces high cost and complexity challenges. Here, we demonstrate the use of a topological material to enhance MRI signal reception. Designed with a stack of weak couplings, this material forms quasi-two-dimensional dual topological boundary states. High properties are achieved through low-loss signal transmission via these topological states, as well as only enhanced local magnetic fields and increased number of channels. Initial tests demonstrate superior performance and accessibility compared to commercial coils, suggesting significant potential. This concept introduces a transformative paradigm for all MRI coil designs.

physics.app-ph

High-Tc superconductivity above 130 K in cubic MH4 compounds at ambient pressure

Hydrides have long been considered promising candidates for achieving room-temperature superconductivity; however, the extremely high pressures typically required for high critical temperatures remain a major challenge in experiment. Here, we propose a class of high-Tc ambient-pressure superconductors with MH4 stoichiometry. These hydrogen-based compounds adopt the bcc PtHg4 structure type, in which hydrogen atoms occupy the one-quarter body-diagonal sites of metal lattices, with the metal atoms acting as chemical templates for hydrogen assembly. Through comprehensive first-principles calculations, we identify three promising superconductors, PtH4, AuH4 and PdH4, with superconducting critical temperatures of 84 K, 89 K, and 133 K, respectively, all surpassing the liquid-nitrogen temperature threshold of 77 K. The remarkable superconducting properties originate from strong electron-phonon coupling associated with hydrogen vibrations, which in turn arise from phonon softening in the mid-frequency range. Our results provide crucial insights into the design of high-Tc superconductors suitable for future experiments and applications at ambient pressure.

cond-mat.supr-con

LLMs Can Also Do Well! Breaking Barriers in Semantic Role Labeling via Large Language Models

Semantic role labeling (SRL) is a crucial task of natural language processing (NLP). Although generative decoder-based large language models (LLMs) have achieved remarkable success across various NLP tasks, they still lag behind state-of-the-art encoder-decoder (BERT-like) models in SRL. In this work, we seek to bridge this gap by equipping LLMs for SRL with two mechanisms: (a) retrieval-augmented generation and (b) self-correction. The first mechanism enables LLMs to leverage external linguistic knowledge such as predicate and argument structure descriptions, while the second allows LLMs to identify and correct inconsistent SRL outputs. We conduct extensive experiments on three widely-used benchmarks of SRL (CPB1.0, CoNLL-2009, and CoNLL-2012). Results demonstrate that our method achieves state-of-the-art performance in both Chinese and English, marking the first successful application of LLMs to surpass encoder-decoder approaches in SRL.

cs.CL

UniSymNet: A Unified Symbolic Network Guided by Transformer

Symbolic Regression (SR) is a powerful technique for automatically discovering mathematical expressions from input data. Mainstream SR algorithms search for the optimal symbolic tree in a vast function space, but the increasing complexity of the tree structure limits their performance. Inspired by neural networks, symbolic networks have emerged as a promising new paradigm. However, most existing symbolic networks still face certain challenges: binary nonlinear operators $\{\times, \div\}$ cannot be naturally extended to multivariate operators, and training with fixed architecture often leads to higher complexity and overfitting. In this work, we propose a Unified Symbolic Network that unifies nonlinear binary operators into nested unary operators and define the conditions under which UniSymNet can reduce complexity. Moreover, we pre-train a Transformer model with a novel label encoding method to guide structural selection, and adopt objective-specific optimization strategies to learn the parameters of the symbolic network. UniSymNet shows high fitting accuracy, excellent symbolic solution rate, and relatively low expression complexity, achieving competitive performance on low-dimensional Standard Benchmarks and high-dimensional SRBench.

cs.LG

Mechanical Amorphization of Glass-Forming Systems Induced by Oscillatory Deformation: The Energy Absorption and Efficiency Control

The kinetic process of mechanical amorphization plays a central role in tailoring material properties. Therefore, a quantitative understanding of how this process depends on loading parameters is critical for optimizing mechanical amorphization and tuning material performance. In this study, we employ molecular dynamics simulations to investigate oscillatory deformation-induced amorphization in three glass-forming intermetallic systems, addressing two unresolved challenges: (1) the relationship between amorphization efficiency and mechanical loading, and (2) energy absorption dynamics during crystal-to-amorphous (CTA) transitions. Our results demonstrate a decoupling between amorphization efficiency--governed by work rate and described by an effective temperature model--and energy absorption, which adheres to the Herschel-Bulkley constitutive relation. Crucially, the melting enthalpy emerges as a key determinant of the energy barrier, establishing a thermodynamic analogy between mechanical amorphization and thermally induced melting. This relationship provides a universally applicable metric to quantify amorphization kinetics. By unifying material properties and loading conditions, this work establishes a predictive framework for controlling amorphization processes. These findings advance the fundamental understanding of deformation-driven phase transitions and offer practical guidelines for designing materials with tailored properties for ultrafast fabrication, ball milling, and advanced mechanical processing techniques.

cond-mat.mtrl-sci

ViSymRe: Vision Multimodal Symbolic Regression

Extracting interpretable equations from observational datasets to describe complex natural phenomena is one of the core goals of artificial intelligence. This field is known as symbolic regression (SR). In recent years, Transformer-based paradigms have become a new trend in SR, addressing the well-known problem of inefficient search. However, the modal heterogeneity between datasets and equations often hinders the convergence and generalization of these models. In this paper, we propose ViSymRe, a Vision Symbolic Regression framework, to explore the positive role of visual modality in enhancing the performance of Transformer-based SR paradigms. To overcome the challenge where the visual SR model is untrainable in high-dimensional scenarios, we present Multi-View Random Slicing (MVRS). By projecting multivariate equations into 2-D space using random affine transformations, MVRS avoids common defects in high-dimensional visualization, such as variable degradation, non-linear interaction missing, and exponentially increasing sampling complexity, enabling ViSymRe to be trained with low computational costs. To support dataset-only deployment of ViSymRe, we design a dual-vision pipeline architecture based on generative techniques, which reconstructs visual features directly from the datasets via an auxiliary Visual Decoder and automatically suppresses the attention weights of reconstruction noise through a proposed Biased Cross-Attention feature fusion module, ensuring that subsequent processes are not affected by noisy modalities. Ablation studies demonstrate the positive contribution of visual modality to improving model convergence level and enhancing various SR metrics. Furthermore, evaluation results on mainstream benchmarks indicate that ViSymRe achieves competitive performance compared to baselines, particularly in low-complexity and rapid-inference scenarios.

cs.LG

A New Adaptive Balanced Augmented Lagrangian Method with Application to ISAC Beamforming Design

In this paper, we consider a class of convex programming problems with linear equality constraints, which finds broad applications in machine learning and signal processing. We propose a new adaptive balanced augmented Lagrangian (ABAL) method for solving these problems. The proposed ABAL method adaptively selects the stepsize parameter and enjoys a low per-iteration complexity, involving only the computation of a proximal mapping of the objective function and the solution of a linear equation. These features make the proposed method well-suited to large-scale problems. We then custom-apply the ABAL method to solve the ISAC beamforming design problem, which is formulated as a nonlinear semidefinite program in a previous work. This customized application requires careful exploitation of the problem's special structure such as the property that all of its signal-to-interference-and-noise-ratio (SINR) constraints hold with equality at the solution and an efficient computation of the proximal mapping of the objective function. Simulation results demonstrate the efficiency of the proposed ABAL method.

eess.SP

Sulfur and sulfur-oxide compounds as potential optically active defects on SWCNTs

Semiconducting single-walled carbon nanotubes (SWCNT) functionalized with covalent defects are a promising class of optoelectronic materials with strong, tunable photoluminescence and demonstrated single photon emission (SPE). Here, we investigate sulfur-oxide containing compounds as a new class of optically active dopants on (6,5) SWCNT. Experimentally, it has been found that when the SWCNT is exposed to sodium dithionite, the resulting compound displays a red-shifted and bright photoluminescence peak that is characteristic of doping with covalent defects. We perform density functional theory calculations on the possible adsorbed compounds that may be the source of doping (S, SO, SO2 and SO3). We predict that the two smallest molecules strongly bind to the SWCNT with binding energies of ~ 1.5-1.8 eV and 0.56 eV for S and SO, respectively, and introduce in-gap electronic states into the bandstructure of the tube consistent with the measured red-shift of (0.1-0.3) eV, consistent with measurements. In contrast, the larger compounds are found to be either unbound or weakly physisorbed with no appreciable impact on the electronic structure of the tube, indicating that they are unlikely to occur. Overall, our study suggests that sulfur-based compounds are promising new dopants for (6,5) SWCNT with tunable electronic properties.

cond-mat.mtrl-sci

Stress-tunable abilities of glass forming and mechanical amorphization

Mechanical amorphization, a widely observed phenomenon, has been utilized to synthesize novel phases by inducing disorder through external loading, thereby expanding the realm of glass-forming systems. Empirically, it has been plausible that mechanical amorphization ability consistently correlates with glass-forming ability. However, through a comprehensive investigation in binary, ternary, and quaternary systems combining neutron diffraction, calorimetric experimental approaches and molecular dynamics simulation, we demonstrate that this impression is only partly true and we reveal that the mechanical amorphization ability can be inversely correlated with the glass forming ability in certain cases To provide insights into these intriguing findings, we present a stress-dependent nucleation theory that offers a coherent explanation for both experimental and simulation results. Our study identifies the intensity of mechanical work, contributed by external stress, as the key control parameter for mechanical amorphization, rendering the ability to tune this process. This discovery not only unravels the underlying correlation between mechanical amorphization and glass-forming ability but also provides a pathway for the design and discovery of new amorphous phases with tailored properties.

cond-mat.mtrl-sci

Research on the Quantum confinement of Carriers in the Type-I Quantum Wells Structure

Quantum confinement is recognized to be an inherent property in low-dimensional structures. Traditionally it is believed that the carriers trapped within the well cannot escape due to the discrete energy levels. However, our previous research has revealed efficient carrier escape in low-dimensional structures, contradicting this conventional understanding. In this study, we review the energy band structure of quantum wells considering it as a superposition of the bulk material dispersion and quantization energy dispersion resulting from the quantum confinement across the whole Brillouin zone. By accounting for all wave vectors, we obtain a certain distribution of carrier energy at each quantization energy level, giving rise to the energy subbands. These results enable carriers to escape from the well under the influence of an electric field. Additionally, we have compiled a comprehensive summary of various energy band scenarios in quantum well structures, relevant to carrier transport. Such a new interpretation holds significant value in deepening our comprehension of low-dimensional energy bands, discovering new physical phenomena, and designing novel devices with superior performance.

cond-mat.mes-hall

Intersubband Transitions in Lead Halide Perovskite-Based Quantum Wells for Mid-Infrared Detectors

Due to their excellent optical and electrical properties as well as versatile growth and fabrication processes, lead halide perovskites have been widely considered as promising candidates for green energy and opto-electronic related applications. Here, we investigate their potential applications at infrared wavelengths by modeling the intersubband transitions in lead halide perovskite-based quantum well systems. Both single-well and double-well structures are studied and their energy levels as well as the corresponding wavefunctions and intersubband transition energies are calculated by solving the one-dimensional Schr\"odinger equations. By adjusting the quantum well and barrier thicknesses, we are able to tune the intersubband transition energies to cover a broad range of infrared wavelengths. We also find that the lead-halide perovskite-based quantum wells possess high absorption coefficients, which are beneficial for their potential applications in infrared photodetectors. The widely tunable transition energies and high absorption coefficients of the perovskite-based quantum well systems, combined with their unique material and electrical properties, may enable an alternative material system for the development of infrared photodetectors.

cond-mat.mtrl-sci

Demonstration of an AI-driven workflow for autonomous high-resolution scanning microscopy

With the continuing advances in scientific instrumentation, scanning microscopes are now able to image physical systems with up to sub-atomic-level spatial resolutions and sub-picosecond time resolutions. Commensurately, they are generating ever-increasing volumes of data, storing and analysis of which is becoming an increasingly difficult prospect. One approach to address this challenge is through self-driving experimentation techniques that can actively analyze the data being collected and use this information to make on-the-fly measurement choices, such that the data collected is sparse but representative of the sample and sufficiently informative. Here, we report the Fast Autonomous Scanning Toolkit (FAST) that combines a trained neural network, a route optimization technique, and efficient hardware control methods to enable a self-driving scanning microscopy experiment. The key features of our method are that: it does not require any prior information about the sample, it has a very low computational cost, and that it uses generic hardware controls with minimal experiment-specific wrapping. We test this toolkit in numerical experiments and a scanning dark-field x-ray microscopy experiment of a $WSe_2$ thin film, where our experiments show that a FAST scan of <25% of the sample is sufficient to produce both a high-fidelity image and a quantitative analysis of the surface distortions in the sample. We show that FAST can autonomously identify all features of interest in the sample while significantly reducing the scan time, the volume of data acquired, and dose on the sample. The FAST toolkit is easy to apply for any scanning microscopy modalities and we anticipate adoption of this technique will empower broader multi-level studies of the evolution of physical phenomena with respect to time, temperature, or other experimental parameters.

physics.app-ph

A semantic hierarchical graph neural network for text classification

The key to the text classification task is language representation and important information extraction, and there are many related studies. In recent years, the research on graph neural network (GNN) in text classification has gradually emerged and shown its advantages, but the existing models mainly focus on directly inputting words as graph nodes into the GNN models ignoring the different levels of semantic structure information in the samples. To address the issue, we propose a new hierarchical graph neural network (HieGNN) which extracts corresponding information from word-level, sentence-level and document-level respectively. Experimental results on several benchmark datasets achieve better or similar results compared to several baseline methods, which demonstrate that our model is able to obtain more useful information for classification from samples.

cs.CL