SearcharxivSearch

arXiv subjects

Shuhua Zhang

Publications and source records attributed to Shuhua Zhang.

12 recordsLinked to original sources

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for example, when one raises their voice or speaks faster during conversation. To address this, we propose a style-controllable speech generation model that augments speech across diverse styles while preserving the target speaker's identity. The proposed system starts with diarized segments from a conventional diarizer. For each diarized segment, it generates augmented speech samples enriched with phonetic and stylistic diversity. And then, speaker embeddings from both the original and generated audio are blended to enhance the system's robustness in grouping segments with high intrinsic intra-speaker variability. We validate our approach on a simulated emotional speech dataset and the truncated AMI dataset, demonstrating significant improvements, with error rate reductions of 49% and 35% on each dataset, respectively.

eess.AS

The Exploratory Multi-Asset Mean-Variance Portfolio Selection using Reinforcement Learning

In this paper, we study the continuous-time multi-asset mean-variance (MV) portfolio selection using a reinforcement learning (RL) algorithm, specifically the soft actor-critic (SAC) algorithm, in the time-varying financial market. A family of Gaussian portfolio selections is derived, and a policy iteration process is crafted to learn the optimal exploratory portfolio selection. We prove the convergence of the policy iteration process theoretically, based on which the SAC algorithm is developed. To improve the algorithm's stability and the learning accuracy in the multi-asset scenario, we divide the model parameters that influence the optimal portfolio selection into three parts, and learn each part progressively. Numerical studies in the simulated and real financial markets confirm the superior performance of the proposed SAC algorithm under various criteria.

q-fin.MF

The mean-variance portfolio selection based on the average and current profitability of the risky asset

We study the continuous-time pre-commitment mean-variance portfolio selection in a time-varying financial market. By introducing two indexes which respectively express the average profitability of the risky asset (AP) and the current profitability of the risky asset (CP), the optimal portfolio selection is represented by AP and CP. Furthermore, instead of the traditional maximum likelihood estimation (MLE) of return rate and volatility of the risky asset, we estimate AP and CP with the second-order variation of an auxiliary wealth process. We prove that the estimations of AP and CP in this paper are more accurate than that in MLE. And, the portfolio selection is implemented in various simulated and real financial markets. Numerical studies confirm the superior performance of our portfolio selection with the estimation of AP and CP under various evaluation criteria.

q-fin.MF

Deep learning radiomics for assessment of gastroesophageal varices in people with compensated advanced chronic liver disease

Objective: Bleeding from gastroesophageal varices (GEV) is a medical emergency associated with high mortality. We aim to construct an artificial intelligence-based model of two-dimensional shear wave elastography (2D-SWE) of the liver and spleen to precisely assess the risk of GEV and high-risk gastroesophageal varices (HRV). Design: A prospective multicenter study was conducted in patients with compensated advanced chronic liver disease. 305 patients were enrolled from 12 hospitals, and finally 265 patients were included, with 1136 liver stiffness measurement (LSM) images and 1042 spleen stiffness measurement (SSM) images generated by 2D-SWE. We leveraged deep learning methods to uncover associations between image features and patient risk, and thus conducted models to predict GEV and HRV. Results: A multi-modality Deep Learning Risk Prediction model (DLRP) was constructed to assess GEV and HRV, based on LSM and SSM images, and clinical information. Validation analysis revealed that the AUCs of DLRP were 0.91 for GEV (95% CI 0.90 to 0.93, p < 0.05) and 0.88 for HRV (95% CI 0.86 to 0.89, p < 0.01), which were significantly and robustly better than canonical risk indicators, including the value of LSM and SSM. Moreover, DLPR was better than the model using individual parameters, including LSM and SSM images. In HRV prediction, the 2D-SWE images of SSM outperform LSM (p < 0.01). Conclusion: DLRP shows excellent performance in predicting GEV and HRV over canonical risk indicators LSM and SSM. Additionally, the 2D-SWE images of SSM provided more information for better accuracy in predicting HRV than the LSM.

q-bio.TO

Application of Knowledge Distillation to Multi-task Speech Representation Learning

Model architectures such as wav2vec 2.0 and HuBERT have been proposed to learn speech representations from audio waveforms in a self-supervised manner. When they are combined with downstream tasks such as keyword spotting and speaker verification, they provide state-of-the-art performance. However, these models use a large number of parameters, the smallest version of which has 95 million parameters. This constitutes a challenge for edge AI device deployments. In this paper, we investigate the application of knowledge distillation to speech representation learning (SRL) models followed by joint fine-tuning with multiple downstream voice-activated tasks. In our experiments on two such tasks, our approach results in nearly 75% reduction in model size while suffering only 0.1% accuracy and 0.9% equal error rate degradation compared to the full-size model. In addition, we show that fine-tuning the SRL models results in a significant performance boost compared to using frozen SRL models.

eess.AS

Multi-task Voice Activated Framework using Self-supervised Learning

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are learned without any task-specific supervision, they can also be useful for other voice-activated tasks like speaker verification, keyword spotting, emotion classification etc. In our work, we propose a general purpose framework for adapting a pre-trained wav2vec 2.0 model for different voice-activated tasks. We develop downstream network architectures that operate on the contextualized speech representations of wav2vec 2.0 to adapt the representations for solving a given task. Finally, we extend our framework to perform multi-task learning by jointly optimizing the network parameters on multiple voice activated tasks using a shared transformer backbone. Both of our single and multi-task frameworks achieve state-of-the-art results in speaker verification and keyword spotting benchmarks. Our best performing models achieve 1.98% and 3.15% EER on VoxCeleb1 test set when trained on VoxCeleb2 and VoxCeleb1 respectively, and 98.23% accuracy on Google Speech Commands v1.0 keyword spotting dataset.

eess.AS

Geo-information system of spread of tuberculosis based on inversion and prediction

The monitoring, analysis and prediction of epidemic spread in the region require the construction of mathematical model, big data processing and visualization because the amount of population and the size of the region could be huge. One of the important steps is refinement of mathematical model, i.e. determination of initial data and coefficients of system of differential equations which describe the epidemiology processes. We analyze numerical method for solving inverse problem of epidemiology based on genetic algorithm and traditional optimization ideas. Numerical results are applied to analysis and prediction of epidemic situation in regions of Russian Federation, Republic of Kazakhstan and People's Republic of China. Due to a great amount of data we use a special Geo-information system for visualization of epidemic process, i.e. a special software named Digital Earth.

q-bio.PE

Differential evolution algorithm of solving an inverse problem for the spatial Solow mathematical model

The differential evolution algorithm is applied to solve the optimization problem to reconstruct the production function (inverse problem) for the spatial Solow mathematical model using additional measurements of the gross domestic product for the fixed points. Since the inverse problem is ill-posed the regularized differential evolution is applied. For getting the optimized solution of the inverse problem the differential evolution algorithm is paralleled to 32 kernels. Numerical results for different technological levels and errors in measured data are presented and discussed.

math.OC

Prediction of Daily PM2.5 Concentration in China Using Data-Driven Ordinary Differential Equations

Accurate reporting and forecasting of PM2.5 concentration are important for improving public health. In this paper, we propose a daily prediction method of PM2.5 concentration by using data-driven ordinary differential equation (ODE) models. Specifically, based on the historical PM2.5 concentration, this method combines genetic programming and orthogonal least square method to evolve the ODE models, which describe the transport of PM2.5 and then uses the data-driven ODEs to predict the air quality in the future. Experiment results show that the ODE models obtain similar prediction results as the typical statistical model, and the prediction results from this method are relatively good. To our knowledge, this is the first attempt to evolve data-driven ODE models to study PM2.5 prediction.

physics.ao-ph

Character Distributions of Classical Chinese Literary Texts: Zipf's Law, Genres, and Epochs

We collect 14 representative corpora for major periods in Chinese history in this study. These corpora include poetic works produced in several dynasties, novels of the Ming and Qing dynasties, and essays and news reports written in modern Chinese. The time span of these corpora ranges between 1046 BCE and 2007 CE. We analyze their character and word distributions from the viewpoint of the Zipf's law, and look for factors that affect the deviations and similarities between their Zipfian curves. Genres and epochs demonstrated their influences in our analyses. Specifically, the character distributions for poetic works of between 618 CE and 1644 CE exhibit striking similarity. In addition, although texts of the same dynasty may tend to use the same set of characters, their character distributions still deviate from each other.

cs.CL

A Stationary Accumulated Projection Method for Linear System of Equations

It is shown in this paper that, almost all current prevalent iterative \mbox{methods} for solving linear system of equations can be classified as what we called extended Krylov subspace methods. In this paper a new type of iterative methods are introduced which do not depend on any Krylov subspaces. This type of methods are based on the so-called accumulated projection technique proposed by authors. It overcomes some shortcomings of classical Row-Projection technique and takes full advantages of the linear system. Comparing with traditional Krylov subspace methods which always depend on the matrix-vector multiplication with some fixed matrix, the newly introduced method (SAP) uses different projection matrices which differ in each step in the iteration process to form an approximate solution. More importantly some particular accelerative schemes (named as MSAP1 and MSAP2) are introduced to improve the convergence of the SAP method. Numerical experiments show some surprisingly improved convergence behavior; some superior experimental behavior of MSAP methods over GMRES and block-Jacobi are demonstrated in some situations.

math.NA

Orthogonally Accumulated Projection Methods for Linear System of Equations

A type of iterative orthogonally accumulated projection methods for solving linear system of equations are proposed in this paper. This type of methods are applications of accumulated projection(AP) technique proposed recently by authors. Instead of searching projections in a sequence of subspaces as done in the original AP approach, these methods try to efficiently construct a sequence of orthonormal vectors while the inner-product between the solution to the system and each vector in the sequence can be easily calculated, thus the solution can be retrieved in finite number of iterations in case of exact arithmetic operations. We also discuss the strategies to handle loss-of-orthogonality during the process of constructing orthonormal vectors. Numerical experiments are provided to demonstrate the efficiency of these methods.

math.NA