SearcharxivSearch

arXiv subjects

Yuzhong Wu

Publications and source records attributed to Yuzhong Wu.

18 recordsLinked to original sources

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

cs.CL

SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not an explicit constraint-search additionally requires the rewrite to remain faithful to the user's stated query intent. Transplanted directly, these models learn a shortcut we term the generic-word dominance effect: they favor generic rewrites that score well on paths but drift from query intent. To address this, we propose SPEAR (Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval), which integrates three components that each target one failure mode: (1) a dual-embedding backbone with auxiliary loss and gradient isolation that shields recall-side semantics from being eroded by CTR-driven ranking signals; (2) a multiplicative gating aggregator that lets a rewrite score high only when both its confidence and item relevance are strong, eliminating the generic-word shortcut; (3) a Dynamic Rewrite Selector that jointly generates request-specific rewrite weights and user-query-conditioned scale and bias terms, allowing both rewrite preference and relevance calibration to adapt to each request. Offline evaluation on 100K held-out industrial search sessions shows that the proposed framework improves rewrite semantic similarity@10 by +18.2 and click recall@10 by +99.5 over the production baseline. In online A/B testing, SPEAR achieves +0.259 in query-view CTR and +0.733 in average reading depth, confirming that improved rewrite selection translates into stronger retrieval and deeper user engagement. The proposed SPEAR system has been fully deployed in Dewu's community search platform since 2025. Our code is available at https://github.com/mallocagi1-cell/spear.

cs.IR

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in isolation, lacking a unified benchmark for domain terminology, age variation, dialects, accents, and low-resource languages, particularly across the Middle East and Southeast Asia, representing over one billion under-evaluated speakers. To address this gap, we introduce GigaSpeechBench, a comprehensive multilingual and multidimensional in-the-wild ASR & AST benchmark comprising 680 hours of human-annotated speech. It features five modules: (1) 12 low-resource Middle Eastern and Southeast Asian languages, plus challenging Japanese and Korean; (2) 6 Chinese dialects; (3) 6 English accents; (4) dense terminology across 12 vertical domains for Chinese and English; and (5) older adult and child speech. We further provide human-annotated Chinese and English translations for 11 languages to support AST evaluation. Extensive evaluations of leading foundation models and commercial APIs reveal significant performance degradation in these challenging settings, exposing critical evaluation blind spots.

eess.AS

The size-velocity dispersion relationship of Galactic HII regions

The size-velocity dispersion ($σ$) relation, while well established for giant HII regions, remains uncertain for their smaller counterparts (physical radii R < 20 pc). Thanks to the LAMOST MRS-N dataset's large sky coverage and high spatial/spectral resolution, we examined this relationship using 10 isolated Galactic HII regions with R < 20 pc. Our results reveal two key findings: (1) these small-size HII regions remarkably follow the same size-$σ$ relation as giant HII regions, suggesting this correlation could serve as a novel distance indicator for Galactic HII regions; and (2) we find distinct dynamical behaviors between younger and older HII regions. Specifically, in younger (< 0.5 Myr), ionization-bounded HII regions, the velocity dispersion shows no correlation with expansion velocity, indicating that turbulence is driven primarily by stellar winds and ionization processes. In contrast, in older (> 0.5 Myr), matter-bounded HII regions, a clear correlation emerges, implying that expansion-driven processes begin to play a significant role in generating turbulence. We therefore propose an evolutionary transition in the primary turbulence mechanisms, from being dominated by stellar winds and radiation to being increasingly influenced by expansion-driven dynamics, during the evolution of HII regions. Considering the small sample size used in this work, particularly the inclusion of only two young HII regions, which also have large uncertainties in their expansion velocities, further confirmation of this interpretation will require higher-resolution 2D spectroscopy to resolve blended kinematic components along the line of sight for more accurate estimation of expansion velocities, along with an expanded sample that specifically includes more young HII regions.

astro-ph.GA

Fun-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep integration with large language models (LLMs). However, LLMs are prone to hallucination, which can significantly degrade user experience in real-world ASR applications. In this paper, we present Fun-ASR, a large-scale, LLM-based ASR system that synergistically combines massive data, large model capacity, LLM integration, and reinforcement learning to achieve state-of-the-art performance across diverse and complex speech recognition scenarios. Moreover, Fun-ASR is specifically optimized for practical deployment, with enhancements in streaming capability, noise robustness, code-switching, hotword customization, and satisfying other real-world application requirements. Experimental results show that while most LLM-based ASR systems achieve strong performance on open-source benchmarks, they often underperform on real industry evaluation sets. Thanks to production-oriented optimizations, Fun-ASR achieves state-of-the-art performance on real application datasets, demonstrating its effectiveness and robustness in practical settings. The code and models are accessible at https://github.com/FunAudioLLM/Fun-ASR .

cs.CL

Spatially-resolved Galactic HII regions observed by LAMOST medium-Resolution Spectroscopic Survey of Nebulae (LAMOST MRS-N)

We present spatially-resolved spectroscopic observations of 10 isolated Galactic HII regions using data from the LAMOST Medium-Resolution Spectroscopic Survey of Nebulae (LAMOST MRS-N). The high spatial resolution of the data allows us to investigate the 1D radial profiles of emission line fluxes (Ha, [S II] and [N II]), flux ratios ([N II]/Ha, [S II]/Ha and [S II]/[N II]), and radial velocities of these three emission lines. Among these regions, two are ionization-bounded, while the remaining eight are matter-bounded. The matter-bounded HII regions exhibit shallower slopes in their radial flux profiles compared to the ionization-bounded ones. In most cases, the [N II]/Ha and [S II]/Ha ratios increase with distance from the center of the HII regions, consistent with model predictions that low-ionization emissions dominate the outer zones of these regions. The two ionization-bounded HII regions have kinematic ages (t) of 0.2 and 0.3 Myr, while the matter-bounded regions span ages from 1 to 12 Myr. For the matter-bounded HII regions, the optical emission flux decreases continuously beyond the photodissociation region (PDR), extending to approximately 1-4 times the radius of the PDR (r_PDR). The escape fraction f_esc of ionizing photons, derived from 1D Ha radial flux profiles, is ~ 0% for ionization-bounded HII regions, while it ranges from 50% to 90% for the matter-bounded HII regions. The correlation between f_esc and t suggests that evolved HII regions (with t > 1 Myr) contribute more significantly to ionizing the surrounding diffuse ionized gas compared to younger, newly formed HII regions.

astro-ph.GA

Diffuse Ionized Gas in the Anti-center of the Milky Way

Using data from the LAMOST Medium-Resolution Spectroscopic Survey of Nebulae, we create a sample of 17,821 diffuse ionized gas (DIG) spectra in the anti-center region of the Milky Way, by excluding fibers in the directions of H II regions and supernova remnants. We then analyze the radial and vertical distributions of three line ratios ([N II]/H$α$, [S II]/H$α$, and [S II]/[N II]), as well as the oxygen abundance. [N II]/H$α$ and [S II]/H$α$ do not exhibit a consistent, monotonic decrease with increasing Galactocentric distance (R$_{gal}$). Instead, they show enhancement within the interarm region, positioned between the Local Arm and the Perseus Arm. [S II]/[N II] has a radial gradient of 0.1415 $\pm$ 0.0646 kpc$^{-1}$ for the inner disk (8.34 $ < R_{gal} < $ 9.65 kpc), and remains nearly flat for the outer disk ($R_{gal} > $ 9.65 kpc). In the vertical direction, [N II]/H$α$, [S II]/H$α$, and [S II]/[N II] increase with increasing Galactic disk height ($|z|$) in both southern and northern disks. Based on the N2S2H$α$ method, which combines [S II]/[N II] and [N II]/H$α$, we estimate the oxygen abundance. The oxygen abundance exhibits a consistent radial gradient with R$_{gal}$, featuring a slope of -0.0559 $\pm$ 0.0209 dex kpc$^{-1}$ for the inner disk and a similar slope of -0.0429 $\pm$ 0.0599 dex kpc$^{-1}$ for the outer disk. A single linear fitting to the entire disk yields a slope of -0.0317 $\pm$ 0.0124 dex kpc$^{-1}$. In the vertical direction, the oxygen abundance decreases with increasing $|z|$ in both southern and northern disks.

astro-ph.GA

Hybrid AHS: A Hybrid of Kalman Filter and Deep Learning for Acoustic Howling Suppression

Deep learning has been recently introduced for efficient acoustic howling suppression (AHS). However, the recurrent nature of howling creates a mismatch between offline training and streaming inference, limiting the quality of enhanced speech. To address this limitation, we propose a hybrid method that combines a Kalman filter with a self-attentive recurrent neural network (SARNN) to leverage their respective advantages for robust AHS. During offline training, a pre-processed signal obtained from the Kalman filter and an ideal microphone signal generated via teacher-forced training strategy are used to train the deep neural network (DNN). During streaming inference, the DNN's parameters are fixed while its output serves as a reference signal for updating the Kalman filter. Evaluation in both offline and streaming inference scenarios using simulated and real-recorded data shows that the proposed method efficiently suppresses howling and consistently outperforms baselines.

eess.AS

A study on joint modeling and data augmentation of multi-modalities for audio-visual scene classification

In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-trained networks trained only on image data sets to extract video embedding; whereas for audio embedding models, we decide to train them from scratch. We explore different neural network architectures for joint modeling to effectively combine the video and audio modalities. Moreover, data augmentation strategies are investigated to increase audio-visual training set size. For the video modality the effectiveness of several operations in RandAugment is verified. An audio-video joint mixup scheme is proposed to further improve AVSC performances. Evaluated on the development set of TAU Urban Audio Visual Scenes 2021, our final system can achieve the best accuracy of 94.2% among all single AVSC systems submitted to DCASE 2021 Task 1b.

cs.MM

A Lottery Ticket Hypothesis Framework for Low-Complexity Device-Robust Neural Acoustic Scene Classification

We propose a novel neural model compression strategy combining data augmentation, knowledge transfer, pruning, and quantization for device-robust acoustic scene classification (ASC). Specifically, we tackle the ASC task in a low-resource environment leveraging a recently proposed advanced neural network pruning mechanism, namely Lottery Ticket Hypothesis (LTH), to find a sub-network neural model associated with a small amount non-zero model parameters. The effectiveness of LTH for low-complexity acoustic modeling is assessed by investigating various data augmentation and compression schemes, and we report an efficient joint framework for low-complexity multi-device ASC, called \emph{Acoustic Lottery}. Acoustic Lottery could compress an ASC model up to $1/10^{4}$ and attain a superior performance (validation accuracy of 79.4% and Log loss of 0.64) compared to its not compressed seed model. All results reported in this work are based on a joint effort of four groups, namely GT-USTC-UKE-Tencent, aiming to address the "Low-Complexity Acoustic Scene Classification (ASC) with Multiple Devices" in the DCASE 2021 Challenge Task 1a.

cs.SD

LAMOST Medium-Resolution Spectral Survey of Galactic Nebulae (LAMOST-MRS-N): Subtraction of Geocoronal Halpha Emission

We introduce a method of subtracting geocoronal Halpha emissions from the spectra of LAMOST medium-resolution spectral survey of Galactic nebulae (LAMOST-MRS-N). The flux ratios of the Halpha sky line to the adjacent OH lambda6554 single line do not show a pattern or gradient distribution in a plate. More interestingly, the ratio is well correlated to solar altitude, which is the angle of the sun relative to the Earth's horizon. It is found that the ratio decreases from 0.8 to 0.2 with the decreasing solar altitude from -17 to -73 degree. Based on this relation, which is described by a linear function, we can construct the Halpha sky component and subtract it from the science spectrum. This method has been applied to the LAMOST-MRS-N data, and the contamination level of the Halpha sky to nebula is reduced from 40% to less than 10%. The new generated spectra will significantly improve the accuracy of the classifications and the measurements of physical parameters of Galactic nebulae.

astro-ph.IM

Robust Feature Learning on Long-Duration Sounds for Acoustic Scene Classification

Acoustic scene classification (ASC) aims to identify the type of scene (environment) in which a given audio signal is recorded. The log-mel feature and convolutional neural network (CNN) have recently become the most popular time-frequency (TF) feature representation and classifier in ASC. An audio signal recorded in a scene may include various sounds overlapping in time and frequency. The previous study suggests that separately considering the long-duration sounds and short-duration sounds in CNN may improve ASC accuracy. This study addresses the problem of the generalization ability of acoustic scene classifiers. In practice, acoustic scene signals' characteristics may be affected by various factors, such as the choice of recording devices and the change of recording locations. When an established ASC system predicts scene classes on audios recorded in unseen scenarios, its accuracy may drop significantly. The long-duration sounds not only contain domain-independent acoustic scene information, but also contain channel information determined by the recording conditions, which is prone to over-fitting. For a more robust ASC system, We propose a robust feature learning (RFL) framework to train the CNN. The RFL framework down-weights CNN learning specifically on long-duration sounds. The proposed method is to train an auxiliary classifier with only long-duration sound information as input. The auxiliary classifier is trained with an auxiliary loss function that assigns less learning weight to poorly classified examples than the standard cross-entropy loss. The experimental results show that the proposed RFL framework can obtain a more robust acoustic scene classifier towards unseen devices and cities.

cs.SD

A New Transition Wolf-Rayet WN/C Star in the Milky Way

We report the discovery of a new transition type Wolf-Rayet (WR) WN/C star in the Galaxy. According to its coordinates (R.A., Dec)J2000 = 18h51m39.7s, -05d34m51.1s, and the distance (7.11 kpc away from Earth) inferred from the second Gaia, data release, it's found that WR 121-16 is located in the Far 3 kpc Arm, and it is 3.75 kpc away from the Galactic Center. The optical spectra obtained by the Large Sky Area Multi-Object Fiber Spectroscopic Telescope (LAMOST) and the 2.16 m telescope, both located at the Xinglong Observatory in China, indicate that this is a WR star of the transitional WN7o/WC subtype. A current stellar mass of about 7.1 M_solar, a mass-loss rate of M_dot = 10^(-4.97) M_solar/yr, a bolometric luminosity of log L/L_solar = 4.88, and a stellar temperature of T_* = 47 kK are derived, by fitting the observed spectrum with a specific Potsdam Wolf-Rayet (PoWR) model. The magnitude in V-band varies between 13.95 and 14.14 mag, while no period is found. Based on the optical spectra, the time domain data, and the indices of the astrometric solution of the Gaia data, WR 121-16 is likely a transitional WN/C single star rather than a WN+WC binary.

astro-ph.SR

An End-to-End Approach to Automatic Speech Assessment for Cantonese-speaking People with Aphasia

Conventional automatic assessment of pathological speech usually follows two main steps: (1) extraction of pathology-specific features; (2) classification or regression on extracted features. Given the great variety of speech and language disorders, feature design is never a straightforward task, and yet it is most crucial to the performance of assessment. This paper presents an end-to-end approach to automatic speech assessment for Cantonese-speaking People With Aphasia (PWA). The assessment is formulated as a binary classification task to discriminate PWA with high scores of subjective assessment from those with low scores. The sequence-to-one Recurrent Neural Network with Gated Recurrent Unit (GRU-RNN) and Convolutional Neural Network (CNN) models are applied to realize the end-to-end mapping from fundamental speech features to the classification result. The pathology-specific features used for assessment can be learned implicitly by the neural network model. Class Activation Mapping (CAM) method is utilized to visualize how those features contribute to the assessment result. Our experimental results show that the end-to-end approach outperforms the conventional two-step approach in the classification task, and confirm that the CNN model is able to learn impairment-related features that are similar to human-designed features. The experimental results also suggest that CNN model performs better than sequence-to-one GRU-RNN model in this specific task.

eess.AS

Enhancing Sound Texture in CNN-Based Acoustic Scene Classification

Acoustic scene classification is the task of identifying the scene from which the audio signal is recorded. Convolutional neural network (CNN) models are widely adopted with proven successes in acoustic scene classification. However, there is little insight on how an audio scene is perceived in CNN, as what have been demonstrated in image recognition research. In the present study, the Class Activation Mapping (CAM) is utilized to analyze how the log-magnitude Mel-scale filter-bank (log-Mel) features of different acoustic scenes are learned in a CNN classifier. It is noted that distinct high-energy time-frequency components of audio signals generally do not correspond to strong activation on CAM, while the background sound texture are well learned in CNN. In order to make the sound texture more salient, we propose to apply the Difference of Gaussian (DoG) and Sobel operator to process the log-Mel features and enhance edge information of the time-frequency image. Experimental results on the DCASE 2017 ASC challenge show that using edge enhanced log-Mel images as input feature of CNN significantly improves the performance of audio scene classification.

cs.SD

Radial velocity measurements from LAMOST medium-resolution spectroscopic observations: A pointing towards the Kepler field

Radial velocity is one of key measurements in understanding the fundamental properties of stars, stellar clusters and the Galaxy. A plate of stars in the Kepler field were observed in May of 2018 with the medium-resolution spectrographs of LAMOST, aiming to test the performance of this new system which is the upgraded equipment of LAMOST after the first five-year regular survey.We present our analysis on the radial velocity measurements (RVs) derived from these data. The results show that slight and significant systematic errors exist among the RVs obtained from the spectra collected by different spectrographs and exposures, respectively. After correcting the systematic errors with different techniques, the precision of RVs reaches ~1.3, ~1.0, ~0.5 and ~0.3 km/s at S/Nr = 10, 20, 50, and 100, respectively. Comparing with the RVs of the standard stars of the APOGEE survey, our RVs are calibrated with a zero-point shift of ~7 km/s. The results indicate that the LAMOST medium-resolution spectroscopic system may provide RVs in a reasonable accuracy and precision for the selected targets.

astro-ph.SR

Reducing Model Complexity for DNN Based Large-Scale Audio Classification

Audio classification is the task of identifying the sound categories that are associated with a given audio signal. This paper presents an investigation on large-scale audio classification based on the recently released AudioSet database. AudioSet comprises 2 millions of audio samples from YouTube, which are human-annotated with 527 sound category labels. Audio classification experiments with the balanced training set and the evaluation set of AudioSet are carried out by applying different types of neural network models. The classification performance and the model complexity of these models are compared and analyzed. While the CNN models show better performance than MLP and RNN, its model complexity is relatively high and undesirable for practical use. We propose two different strategies that aim at constructing low-dimensional embedding feature extractors and hence reducing the number of model parameters. It is shown that the simplified CNN model has only 1/22 model parameters of the original model, with only a slight degradation of performance.

cs.SD

LAMOST Spectroscopic Survey of the Galactic Anticentre (LSS-GAC): the second release of value-added catalogues

We present the second release of value-added catalogues of the LAMOST Spectroscopic Survey of the Galactic Anticentre (LSS-GAC DR2). The catalogues present values of radial velocity $V_{\rm r}$, atmospheric parameters --- effective temperature $T_{\rm eff}$, surface gravity log$g$, metallicity [Fe/H], $α$-element to iron (metal) abundance ratio [$α$/Fe] ([$α$/M]), elemental abundances [C/H] and [N/H], and absolute magnitudes ${\rm M}_V$ and ${\rm M}_{K_{\rm s}}$ deduced from 1.8 million spectra of 1.4 million unique stars targeted by the LSS-GAC since September 2011 until June 2014. The catalogues also give values of interstellar reddening, distance and orbital parameters determined with a variety of techniques, as well as proper motions and multi-band photometry from the far-UV to the mid-IR collected from the literature and various surveys. Accuracies of radial velocities reach 5kms$^{-1}$ for late-type stars, and those of distance estimates range between 10 -- 30 per cent, depending on the spectral signal-to-noise ratios. Precisions of [Fe/H], [C/H] and [N/H] estimates reach 0.1dex, and those of [$α$/Fe] and [$α$/M] reach 0.05dex. The large number of stars, the contiguous sky coverage, the simple yet non-trivial target selection function and the robust estimates of stellar radial velocities and atmospheric parameters, distances and elemental abundances, make the catalogues a valuable data set to study the structure and evolution of the Galaxy, especially the solar-neighbourhood and the outer disk.

astro-ph.GA