Searcharxiv⌕ Search

arXiv subjects

Li Wan

Publications and source records attributed to Li Wan.

At least 37 records · Page 2Linked to original sources

Lightweight Conceptual Dictionary Learning for Text Classification Using Information Compression

We propose a novel, lightweight supervised dictionary learning framework for text classification based on data compression and representation. This two-phase algorithm initially employs the Lempel-Ziv-Welch (LZW) algorithm to construct a dictionary from text datasets, focusing on the conceptual significance of dictionary elements. Subsequently, dictionaries are refined considering label data, optimizing dictionary atoms to enhance discriminative power based on mutual information and class distribution. This process generates discriminative numerical representations, facilitating the training of simple classifiers such as SVMs and neural networks. We evaluate our algorithm's information-theoretic performance using information bottleneck principles and introduce the information plane area rank (IPAR) as a novel metric to quantify the information-theoretic performance. Tested on six benchmark text datasets, our algorithm competes closely with top models, especially in limited-vocabulary contexts, using significantly fewer parameters. \review{Our algorithm closely matches top-performing models, deviating by only ~2\% on limited-vocabulary datasets, using just 10\% of their parameters. However, it falls short on diverse-vocabulary datasets, likely due to the LZW algorithm's constraints with low-repetition data. This contrast highlights its efficiency and limitations across different dataset types.

cs.CL↗

FADI-AEC: Fast Score Based Diffusion Model Guided by Far-end Signal for Acoustic Echo Cancellation

Despite the potential of diffusion models in speech enhancement, their deployment in Acoustic Echo Cancellation (AEC) has been restricted. In this paper, we propose DI-AEC, pioneering a diffusion-based stochastic regeneration approach dedicated to AEC. Further, we propose FADI-AEC, fast score-based diffusion AEC framework to save computational demands, making it favorable for edge devices. It stands out by running the score model once per frame, achieving a significant surge in processing efficiency. Apart from that, we introduce a novel noise generation technique where far-end signals are utilized, incorporating both far-end and near-end signals to refine the score model's accuracy. We test our proposed method on the ICASSP2023 Microsoft deep echo cancellation challenge evaluation dataset, where our method outperforms some of the end-to-end methods and other diffusion based echo cancellation methods.

eess.AS↗

Non-equilibrium ensemble theory for thermal transport in anharmonic crystals

We propose an ensemble theory for the non-equilibrium statistics to study the thermal transport in anharmonic crystals. In the theory, lattice vibrations of the crystals are quantized by local Bosons(LBs), instead of Phonons as usually used for the thermal transport. LBs are driven by the temperature gradient and move from atom to atom in the crystals. Based on the LBs, anharmonic interactions between atoms in the crystals can be fully considered. To demonstrate our theory, we study the thermal transport in an atomic chain with a temperature drop applied on the two ends of the chain. We observe a Rabi-like oscillation in the transport of the LBs, from which we define the thermal current to get the thermal conductivity of the chain. Results show that the thermal conductivity is enhanced slowly with the increasing of the anharmonic interaction and decreases rapidly if the anharmonic interaction is increased further. In the present study, we only focus on the steady state, and the fluctuations of the thermal currents are not considered.

cond-mat.stat-mech↗

Handling the Alignment for Wake Word Detection: A Comparison Between Alignment-Based, Alignment-Free and Hybrid Approaches

Wake word detection exists in most intelligent homes and portable devices. It offers these devices the ability to "wake up" when summoned at a low cost of power and computing. This paper focuses on understanding alignment's role in developing a wake-word system that answers a generic phrase. We discuss three approaches. The first is alignment-based, where the model is trained with frame-wise cross-entropy. The second is alignment-free, where the model is trained with CTC. The third, proposed by us, is a hybrid solution in which the model is trained with a small set of aligned data and then tuned with a sizeable unaligned dataset. We compare the three approaches and evaluate the impact of the different aligned-to-unaligned ratios for hybrid training. Our results show that the alignment-free system performs better than the alignment-based for the target operating point, and with a small fraction of the data (20%), we can train a model that complies with our initial constraints.

cs.CL↗

Anomalous circularly polarized light emission in organic light-emitting diodes caused by orbital-momentum locking

Chiral circularly polarized (CP) light is central to many photonic technologies, from optical communication of spin information to novel display and imaging technologies. As such, there has been significant effort in the development of chiral emissive materials that allow for the emission of strongly dissymmetric CP light from organic light-emitting diodes (OLEDs). A consensus for chiral emission in such devices is that the molecular chirality of the active layer determines the favored light handedness of CP emission, regardless of the light-emitting direction. Here, we discover that, unconventionally, oppositely propagating CP light exhibits opposite handedness, and reversing the current-flow in OLEDs also switches the handedness of the emitted CP light. This direction-dependent CP emission boosts the net polarization rate by orders of magnitude by resolving an established issue in CP-OLEDs, where the CP light reflected by the back electrode typically erodes the measured dissymmetry. Through detailed theoretical analysis, we assign this anomalous CP emission to a ubiquitous topological electronic property in chiral materials, namely the orbital-momentum locking. Our work paves the way to design new chiroptoelectronic devices and probes the close connections between chiral materials, topological electrons, and CP light in the quantum regime.

cond-mat.mtrl-sci↗

LiCo-Net: Linearized Convolution Network for Hardware-efficient Keyword Spotting

This paper proposes a hardware-efficient architecture, Linearized Convolution Network (LiCo-Net) for keyword spotting. It is optimized specifically for low-power processor units like microcontrollers. ML operators exhibit heterogeneous efficiency profiles on power-efficient hardware. Given the exact theoretical computation cost, int8 operators are more computation-effective than float operators, and linear layers are often more efficient than other layers. The proposed LiCo-Net is a dual-phase system that uses the efficient int8 linear operators at the inference phase and applies streaming convolutions at the training phase to maintain a high model capacity. The experimental results show that LiCo-Net outperforms single-value decomposition filter (SVDF) on hardware efficiency with on-par detection performance. Compared to SVDF, LiCo-Net reduces cycles by 40% on HiFi4 DSP.

cs.LG↗

Ergodicity in glass relaxation

We derive an equation for the glass relaxation. In the derivation, the Zwanzig-Mori projection method is not applied explicitly, which makes our equation different from the mode coupling theory. Due to the nonlinearity, it is difficult to solve the equation to get the full behaviors of the glass relaxation. But we can simplify the equation when time approaches infinity and obtain the static result analytically. The static result shows that the density correlation function decays to zero finally, meaning that the glass relaxation is ergodic. In this study, we also find that the force fluctuation of one individual particle averaged in the glass is sensitive to the temperature and is suggested to be a parameter to reflect the structural transition for the glass relaxation.

cond-mat.dis-nn↗

Self-Supervised Speaker Verification with Simple Siamese Network and Self-Supervised Regularization

Training speaker-discriminative and robust speaker verification systems without speaker labels is still challenging and worthwhile to explore. In this study, we propose an effective self-supervised learning framework and a novel regularization strategy to facilitate self-supervised speaker representation learning. Different from contrastive learning-based self-supervised learning methods, the proposed self-supervised regularization (SSReg) focuses exclusively on the similarity between the latent representations of positive data pairs. We also explore the effectiveness of alternative online data augmentation strategies on both the time domain and frequency domain. With our strong online data augmentation strategy, the proposed SSReg shows the potential of self-supervised learning without using negative pairs and it can significantly improve the performance of self-supervised speaker representation learning with a simple Siamese network architecture. Comprehensive experiments on the VoxCeleb datasets demonstrate that our proposed self-supervised approach obtains a 23.4% relative improvement by adding the effective self-supervised regularization and outperforms other previous works.

eess.AS↗

Speaker Diarization with LSTM

For many years, i-vector based audio embedding techniques were the dominant approach for speaker verification and speaker diarization applications. However, mirroring the rise of deep learning in various domains, neural network based audio embeddings, also known as d-vectors, have consistently demonstrated superior speaker verification performance. In this paper, we build on the success of d-vector based speaker verification systems to develop a new d-vector based approach to speaker diarization. Specifically, we combine LSTM-based d-vector audio embeddings with recent work in non-parametric clustering to obtain a state-of-the-art speaker diarization system. Our system is evaluated on three standard public datasets, suggesting that d-vector based diarization systems offer significant advantages over traditional i-vector based systems. We achieved a 12.0% diarization error rate on NIST SRE 2000 CALLHOME, while our model is trained with out-of-domain data from voice search logs.

eess.AS↗

Hierarchical Structural Analysis Method for Complex Equation-oriented Models

Structural analysis is a method for verifying equation-oriented models in the design of industrial systems. Existing structural analysis methods need flattening of the hierarchical models into an equation system for analysis. However, the large-scale equations in complex models make structural analysis difficult. Aimed to address the issue, this study proposes a hierarchical structural analysis method by exploring the relationship between the singularities of the hierarchical equation-oriented model and its components. This method obtains the singularity of a hierarchical equation-oriented model by analyzing a dummy model constructed with the parts from the decomposing results of its components. Based on this, the structural singularity of a complex model can be obtained by layer-by-layer analysis according to their natural hierarchy. The hierarchical structural analysis method can reduce the equation scale in each analysis and achieve efficient structural analysis of very complex models. This method can be adaptively applied to nonlinear-algebraic and differential-algebraic equation models. The main algorithms, application cases and comparison with the existing methods are present in this paper. The complexity analysis results show the enhanced efficiency of the proposed method in the structural analysis of complex equation-oriented models. Compared with the existing methods, the time complexity of the proposed method is improved significantly.

cs.OH↗

Excitations of Atomic Vibrations in Amorphous Solids

We study excitations of atomic vibrations in the reciprocal space for amorphous solids. There are two kinds of excitations we obtained, collective excitation and local excitation. The collective excitation is the collective vibration of atoms in the amorphous solids while the local excitation is stimulated locally by a single atom vibrating in the solids. We introduce a continuous wave vector for the study and transform the equations of atomic vibrations from the real space to the reciprocal space. We take the amorphous silicon as an example and calculate the structures of the excitations in the reciprocal space. Results show that an excitation is a wave packet composed of a collection of plane waves. We also find a periodical structure in the reciprocal space for the collective excitation with longitudinal vibrations, which is originated from the local order of the structure in the real space of the amorphous solid.

cond-mat.dis-nn↗

Generalized End-to-End Loss for Speaker Verification

In this paper, we propose a new loss function called generalized end-to-end (GE2E) loss, which makes the training of speaker verification models more efficient than our previous tuple-based end-to-end (TE2E) loss function. Unlike TE2E, the GE2E loss function updates the network in a way that emphasizes examples that are difficult to verify at each step of the training process. Additionally, the GE2E loss does not require an initial stage of example selection. With these properties, our model with the new loss function decreases speaker verification EER by more than 10%, while reducing the training time by 60% at the same time. We also introduce the MultiReader technique, which allows us to do domain adaptation - training a more accurate model that supports multiple keywords (i.e. "OK Google" and "Hey Google") as well as multiple dialects.

eess.AS↗

Signal Combination for Language Identification

Google's multilingual speech recognition system combines low-level acoustic signals with language-specific recognizer signals to better predict the language of an utterance. This paper presents our experience with different signal combination methods to improve overall language identification accuracy. We compare the performance of a lattice-based ensemble model and a deep neural network model to combine signals from recognizers with that of a baseline that only uses low-level acoustic signals. Experimental results show that the deep neural network model outperforms the lattice-based ensemble model, and it reduced the error rate from 5.5% in the baseline to 4.3%, which is a 21.8% relative reduction.

cs.LG↗

Poisson-Boltzmann Equation with a Random Field for Charged Fluids

The classical Poisson-Boltzmann equation (CPBE), which is a mean field theory by averaging the ion fluctuation, has been widely used to study ion distributions in charged fluids. In this study, we derive a modified Poisson-Boltzmann equation with a random field from the field theory and recover the ion fluctuation through a multiplicative noise added in the CPBE. The Poisson-Boltzmann equation with a random field (RFPBE) captures the effect of the ion fluctuation and gives different ion distributions in the charged fluids compared to the CPBE. To solve the RFPBE, we propose a Monte Carlo method based on the path integral representation. Numerical results show that the effect of the ion fluctuation strengthens the ion diffusion into the domain and intends to distribute the ions in the fluid uniformly. The final ion distribution in the fluid is determined by the competition between the ion fluctuation and the electrostatic forces exerted by the boundaries. The RFPBE is general and feasible for high dimensional systems by taking the advantage of the Monte Carlo method. We use the RFPBE to study a two dimensional system as an example, in which the effect of ion fluctuation is clearly captured.

physics.chem-ph↗

Tuplemax Loss for Language Identification

In many scenarios of a language identification task, the user will specify a small set of languages which he/she can speak instead of a large set of all possible languages. We want to model such prior knowledge into the way we train our neural networks, by replacing the commonly used softmax loss function with a novel loss function named tuplemax loss. As a matter of fact, a typical language identification system launched in North America has about 95% users who could speak no more than two languages. Using the tuplemax loss, our system achieved a 2.33% error rate, which is a relative 39.4% improvement over the 3.85% error rate of standard softmax loss method.

eess.AS↗

Impurity-Induced Environmental Quantum Phase Transitions in the Quadratic-Coupling Spin-Boson Model

We study the zero temperature properties of the sub-Ohmic spin-boson model with quadratic spin-boson coupling. This model describes experimental set ups at the optimal working point where the linear-coupling between the qubit (spin) and the environmental noise (bosons) is zero and the leading coupling is quadratic. In the strong coupling regime, we find that the existence of spin induces quantum phase transitions (QPTs) between two states of environment: the normal state and a state with local distortions. The phase diagram contains both continuous and the first-order QPTs, with non-trivial critical properties obtained exactly. At the QPTs, the equilibrium state spin dynamics bears power-law $ω$ dependence in the small frequency limit and a robust coherent Rabi oscillation at high frequency. We discuss the feasibility of observing such environmental QPTs in the qubit-related experiments.

cond-mat.mes-hall↗

Weak order in averaging principle for stochastic differential equations with jumps

The present article deals with the averaging principle for a two-time-scale system of jump-diffusion stochastic differential equation. Under suitable conditions, the weak error is expanded in powers of timescale parameter. It is proved that the rate of weak convergence to the averaged dynamics is of order $1$. This reveals the rate of weak convergence is essentially twice that of strong convergence.

math.PR↗

Weak order in averaging principle for two-time-scale stochastic partial differential equations

This work is devoted to averaging principle of a two-time-scale stochastic partial differential equation on a bounded interval $[0, l]$, where both the fast and slow components are directly perturbed by additive noises. Under some regular conditions on drift coefficients, it is proved that the rate of weak convergence for the slow variable to the averaged dynamics is of order $1-\varepsilon$ for arbitrarily small $\varepsilon>0$. The proof is based on an asymptotic expansion of solutions to Kolmogorov equations associated with the multiple-time-scale system.

math.PR↗