SearcharxivSearch

arXiv subjects

Mei Yu

Publications and source records attributed to Mei Yu.

At least 19 recordsLinked to original sources

Non-Markovian dynamics of the giant atom beyond the rotating-wave approximation

We study the non-Markovian dynamics of a giant artificial atom coupled to a one-dimensional acoustic waveguide beyond the rotating-wave and weak-coupling approximations. By combining an optimized ESPRIT-based decomposition of the bath correlation function with the hierarchical equations of motion (HEOM), we achieve numerically exact simulations in regimes with long memory times, finite temperature, and strong system-bath coupling. Benchmarking against analytical results reveals the breakdown of perturbative non-Markovian approaches such as Redfield theory even at weak coupling in the presence of delay-induced memory. We further show that non-Markovian features, including excitation revivals, remain robust at finite temperature and can be enhanced by increasing the system-bath coupling strength. Our approach provides a versatile framework for studying non-Markovian quantum dynamics in structured environments relevant to giant-atom platforms.

quant-ph

Existence, Nonexistence, and Symmetry of Positive Solutions for Fractional Laplacian Problems

This paper studies the properties of solutions to a class of elliptic and parabolic problems involving the fractional Laplacian. By applying the mountain pass theorem, we prove the existence of bounded classical positive solutions in the subcritical regime. Moreover, using the method of moving planes, we establish that these solutions are symmetric or monotone in the first variable. In contrast, we show that no such solutions exist in the supercritical or negative exponent cases. An analysis of the asymptotic behavior of solutions at infinity provides further insight into their profiles, which supports applications to real-world problems. The approaches developed in this work can also be extended to a wider range of nonlocal elliptic and parabolic equations, including those with more general operators and nonlinearities.

math.AP

Quantum memory in spontaneous emission processes

Quantum memory effects are essential in understanding and controlling open quantum systems, yet distinguishing them from classical memory remains challenging. We introduce a convex geometric framework to analyze quantum memory propagating in non-Markovian processes. We prove that classical memory between two time points is fundamentally bounded and introduce a robustness measure for quantum memory based on convex geometry. This admits an efficient experimental characterization by linear witnesses of quantum memory, bypassing full process tomography. We prove that any memory effects present in the spontaneous emission process of two- and three-level atomic systems are necessarily quantum, suggesting a pervasive role of quantum memory in quantum optics. Giant artificial atoms are discussed as a readily available test platform.

quant-ph

Focusing Image Generation to Mitigate Spurious Correlations

Instance features in images exhibit spurious correlations with background features, affecting the training process of deep neural classifiers. This leads to insufficient attention to instance features by the classifier, resulting in erroneous classification outcomes. In this paper, we propose a data augmentation method called Spurious Correlations Guided Synthesis (SCGS) that mitigates spurious correlations through image generation model. This approach does not require expensive spurious attribute (group) labels for the training data and can be widely applied to other debiasing methods. Specifically, SCGS first identifies the incorrect attention regions of a pre-trained classifier on the training images, and then uses an image generation model to generate new training data based on these incorrect attended regions. SCGS increases the diversity and scale of the dataset to reduce the impact of spurious correlations on classifiers. Changes in the classifier's attention regions and experimental results on three different domain datasets demonstrate that this method is effective in reducing the classifier's reliance on spurious correlations.

cs.CV

WaterMamba: Visual State Space Model for Underwater Image Enhancement

Underwater imaging often suffers from low quality due to factors affecting light propagation and absorption in water. To improve image quality, some underwater image enhancement (UIE) methods based on convolutional neural networks (CNN) and Transformer have been proposed. However, CNN-based UIE methods are limited in modeling long-range dependencies, and Transformer-based methods involve a large number of parameters and complex self-attention mechanisms, posing efficiency challenges. Considering computational complexity and severe underwater image degradation, a state space model (SSM) with linear computational complexity for UIE, named WaterMamba, is proposed. We propose spatial-channel omnidirectional selective scan (SCOSS) blocks comprising spatial-channel coordinate omnidirectional selective scan (SCCOSS) modules and a multi-scale feedforward network (MSFFN). The SCOSS block models pixel and channel information flow, addressing dependencies. The MSFFN facilitates information flow adjustment and promotes synchronized operations within SCCOSS modules. Extensive experiments showcase WaterMamba's cutting-edge performance with reduced parameters and computational resources, outperforming state-of-the-art methods on various datasets, validating its effectiveness and generalizability. The code will be released on GitHub after acceptance.

cs.CV

Windformer:Bi-Directional Long-Distance Spatio-Temporal Network For Wind Speed Prediction

Wind speed prediction is critical to the management of wind power generation. Due to the large range of wind speed fluctuations and wake effect, there may also be strong correlations between long-distance wind turbines. This difficult-to-extract feature has become a bottleneck for improving accuracy. History and future time information includes the trend of airflow changes, whether this dynamic information can be utilized will also affect the prediction effect. In response to the above problems, this paper proposes Windformer. First, Windformer divides the wind turbine cluster into multiple non-overlapping windows and calculates correlations inside the windows, then shifts the windows partially to provide connectivity between windows, and finally fuses multi-channel features based on detailed and global information. To dynamically model the change process of wind speed, this paper extracts time series in both history and future directions simultaneously. Compared with other current-advanced methods, the Mean Square Error (MSE) of Windformer is reduced by 0.5\% to 15\% on two datasets from NERL.

eess.SP

Criticality-Enhanced Precision in Phase Thermometry

Temperature estimation of interacting quantum many-body systems is both a challenging task and topic of interest in quantum metrology, given that critical behavior at phase transitions can boost the metrological sensitivity. Here we study non-invasive quantum thermometry of a finite, two-dimensional Ising spin lattice based on measuring the non-Markovian dephasing dynamics of a spin probe coupled to the lattice. We demonstrate a strong critical enhancement of the achievable precision in terms of the quantum Fisher information, which depends on the coupling range and the interrogation time. Our numerical simulations are compared to instructive analytic results for the critical scaling of the sensitivity in the Curie-Weiss model of a fully connected lattice and to the mean-field description in the thermodynamic limit, both of which fail to describe the critical spin fluctuations on the lattice the spin probe is sensitive to. Phase metrology could thus help to investigate the critical behaviour of finite many-body systems beyond the validity of mean-field models.

quant-ph

Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly from speech and ignoring the importance of phoneme information which may limit the performance of AAI. To this end, we propose a novel network called SPN that uses two different streams to carry out the AAI task. Firstly, to improve the performance of speaker-independent experiment, we propose a new phoneme stream network to estimate the articulatory parameters as the phoneme features. To the best of our knowledge, this is the first work that extracts the speaker-independent features from phonemes to improve the performance of AAI. Secondly, in order to better represent the speech information, we train a speech stream network to combine the local features and the global features. Compared with state-of-the-art (SOTA), the proposed method reduces 0.18mm on RMSE and increases 6.0% on Pearson correlation coefficient in the speaker-independent experiment. The code has been released at https://github.com/liujinyu123/AAINetwork-SPN.

cs.SD

Exact Entanglement Dynamics of Two Spins in Finite Baths

We consider the buildup and decay of two-spin entanglement through phase interactions in a finite environment of surrounding spins, as realized in quantum computing platforms based on arrays of atoms, molecules, or nitrogen vacancy centers. The non-Markovian dephasing caused by the spin environment through Ising-type phase interactions can be solved exactly and compared to an effective Markovian treatment based on collision models. In a first case study on a dynamic lattice of randomly hopping spins, we find that non-Markovianity boosts the dephasing rate caused by nearest neighbour interactions with the surroundings, degrading the maximum achievable entanglement. However, we also demonstrate that additional three-body interactions can mitigate this degradation, and that randomly timed reset operations performed on the two-spin system can help sustain a finite average amount of steady-state entanglement. In a second case study based on a model nuclear magnetic resonance system, we elucidate the role of bath correlations at finite temperature on non-Markovian dephasing. They speed up the dephasing at low temperatures while slowing it down at high temperatures, compared to an uncorrelated bath, which is related to the number of thermally accessible spin configurations with and without interactions.

quant-ph

MVNet: Memory Assistance and Vocal Reinforcement Network for Speech Enhancement

Speech enhancement improves speech quality and promotes the performance of various downstream tasks. However, most current speech enhancement work was mainly devoted to improving the performance of downstream automatic speech recognition (ASR), only a relatively small amount of work focused on the automatic speaker verification (ASV) task. In this work, we propose a MVNet consisted of a memory assistance module which improves the performance of downstream ASR and a vocal reinforcement module which boosts the performance of ASV. In addition, we design a new loss function to improve speaker vocal similarity. Experimental results on the Libri2mix dataset show that our method outperforms baseline methods in several metrics, including speech quality, intelligibility, and speaker vocal similarity et al.

cs.SD

Reinforced Swin-Convs Transformer for Underwater Image Enhancement

Underwater Image Enhancement (UIE) technology aims to tackle the challenge of restoring the degraded underwater images due to light absorption and scattering. To address problems, a novel U-Net based Reinforced Swin-Convs Transformer for the Underwater Image Enhancement method (URSCT-UIE) is proposed. Specifically, with the deficiency of U-Net based on pure convolutions, we embedded the Swin Transformer into U-Net for improving the ability to capture the global dependency. Then, given the inadequacy of the Swin Transformer capturing the local attention, the reintroduction of convolutions may capture more local attention. Thus, we provide an ingenious manner for the fusion of convolutions and the core attention mechanism to build a Reinforced Swin-Convs Transformer Block (RSCTB) for capturing more local attention, which is reinforced in the channel and the spatial attention of the Swin Transformer. Finally, the experimental results on available datasets demonstrate that the proposed URSCT-UIE achieves state-of-the-art performance compared with other methods in terms of both subjective and objective evaluations. The code will be released on GitHub after acceptance.

cs.CV

Nash Equilibrium Seeking for General Linear Systems with Disturbance Rejection

This paper explores aggregative games in a network of general linear systems subject to external disturbances. To deal with external disturbances, distributed strategy-updating rules based on internal model are proposed for the case with perfect and imperfect information, respectively. Different from existing algorithms based on gradient dynamics, by introducing the integral of gradient of cost functions on the basis of passive theory, the rules are proposed to force the strategies of all players to evolve to Nash equilibrium regardless the effect of disturbances. The convergence of the two strategy-updating rules is analyzed via Lyapunov stability theory, passive theory and singular perturbation theory. Simulations are presented to verify the obtained results.

math.OC

Domain Adaptation for Underwater Image Enhancement

Recently, learning-based algorithms have shown impressive performance in underwater image enhancement. Most of them resort to training on synthetic data and achieve outstanding performance. However, these methods ignore the significant domain gap between the synthetic and real data (i.e., interdomain gap), and thus the models trained on synthetic data often fail to generalize well to real underwater scenarios. Furthermore, the complex and changeable underwater environment also causes a great distribution gap among the real data itself (i.e., intra-domain gap). However, almost no research focuses on this problem and thus their techniques often produce visually unpleasing artifacts and color distortions on various real images. Motivated by these observations, we propose a novel Two-phase Underwater Domain Adaptation network (TUDA) to simultaneously minimize the inter-domain and intra-domain gap. Concretely, a new dual-alignment network is designed in the first phase, including a translation part for enhancing realism of input images, followed by an enhancement part. With performing image-level and feature-level adaptation in two parts by jointly adversarial learning, the network can better build invariance across domains and thus bridge the inter-domain gap. In the second phase, we perform an easy-hard classification of real data according to the assessed quality of enhanced images, where a rank-based underwater quality assessment method is embedded. By leveraging implicit quality information learned from rankings, this method can more accurately assess the perceptual quality of enhanced images. Using pseudo labels from the easy part, an easy-hard adaptation technique is then conducted to effectively decrease the intra-domain gap between easy and hard samples.

cs.CV

Single Underwater Image Enhancement Using an Analysis-Synthesis Network

Most deep models for underwater image enhancement resort to training on synthetic datasets based on underwater image formation models. Although promising performances have been achieved, they are still limited by two problems: (1) existing underwater image synthesis models have an intrinsic limitation, in which the homogeneous ambient light is usually randomly generated and many important dependencies are ignored, and thus the synthesized training data cannot adequately express characteristics of real underwater environments; (2) most of deep models disregard lots of favorable underwater priors and heavily rely on training data, which extensively limits their application ranges. To address these limitations, a new underwater synthetic dataset is first established, in which a revised ambient light synthesis equation is embedded. The revised equation explicitly defines the complex mathematical relationship among intensity values of the ambient light in RGB channels and many dependencies such as surface-object depth, water types, etc, which helps to better simulate real underwater scene appearances. Secondly, a unified framework is proposed, named ANA-SYN, which can effectively enhance underwater images under collaborations of priors (underwater domain knowledge) and data information (underwater distortion distribution). The proposed framework includes an analysis network and a synthesis network, one for priors exploration and another for priors integration. To exploit more accurate priors, the significance of each prior for the input image is explored in the analysis network and an adaptive weighting module is designed to dynamically recalibrate them. Meanwhile, a novel prior guidance module is introduced in the synthesis network, which effectively aggregates the prior and data features and thus provides better hybrid information to perform the more reasonable image enhancement.

cs.CV

An Attention Self-supervised Contrastive Learning based Three-stage Model for Hand Shape Feature Representation in Cued Speech

Cued Speech (CS) is a communication system for deaf people or hearing impaired people, in which a speaker uses it to aid a lipreader in phonetic level by clarifying potentially ambiguous mouth movements with hand shape and positions. Feature extraction of multi-modal CS is a key step in CS recognition. Recent supervised deep learning based methods suffer from noisy CS data annotations especially for hand shape modality. In this work, we first propose a self-supervised contrastive learning method to learn the feature representation of image without using labels. Secondly, a small amount of manually annotated CS data are used to fine-tune the first module. Thirdly, we present a module, which combines Bi-LSTM and self-attention networks to further learn sequential features with temporal and contextual information. Besides, to enlarge the volume and the diversity of the current limited CS datasets, we build a new British English dataset containing 5 native CS speakers. Evaluation results on both French and British English datasets show that our model achieves over 90% accuracy in hand shape recognition. Significant improvements of 8.75% (for French) and 10.09% (for British English) are achieved in CS phoneme recognition correctness compared with the state-of-the-art.

cs.MM

Cross-Modal Knowledge Distillation Method for Automatic Cued Speech Recognition

Cued Speech (CS) is a visual communication system for the deaf or hearing impaired people. It combines lip movements with hand cues to obtain a complete phonetic repertoire. Current deep learning based methods on automatic CS recognition suffer from a common problem, which is the data scarcity. Until now, there are only two public single speaker datasets for French (238 sentences) and British English (97 sentences). In this work, we propose a cross-modal knowledge distillation method with teacher-student structure, which transfers audio speech information to CS to overcome the limited data problem. Firstly, we pretrain a teacher model for CS recognition with a large amount of open source audio speech data, and simultaneously pretrain the feature extractors for lips and hands using CS data. Then, we distill the knowledge from teacher model to the student model with frame-level and sequence-level distillation strategies. Importantly, for frame-level, we exploit multi-task learning to weigh losses automatically, to obtain the balance coefficient. Besides, we establish a five-speaker British English CS dataset for the first time. The proposed method is evaluated on French and British English CS datasets, showing superior CS recognition performance to the state-of-the-art (SOTA) by a large margin.

cs.MM

Liquidation, Leverage and Optimal Margin in Bitcoin Futures Markets

Using the generalized extreme value theory to characterize tail distributions, we address liquidation, leverage, and optimal margins for bitcoin long and short futures positions. The empirical analysis of perpetual bitcoin futures on BitMEX shows that (1) daily forced liquidations to out- standing futures are substantial at 3.51%, and 1.89% for long and short; (2) investors got forced liquidation do trade aggressively with average leverage of 60X; and (3) exchanges should elevate current 1% margin requirement to 33% (3X leverage) for long and 20% (5X leverage) for short to reduce the daily margin call probability to 1%. Our results further suggest normality assumption on return significantly underestimates optimal margins. Policy implications are also discussed.

q-fin.TR

Three-Dimensional Lip Motion Network for Text-Independent Speaker Recognition

Lip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a textdependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip- Motion-Network-for-Text-Independent-Speaker-Recognition.

cs.CV