SearcharxivSearch

arXiv subjects

Bin Ma

Publications and source records attributed to Bin Ma.

At least 19 recordsLinked to original sources

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

cs.CL

TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation

Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range contexts for monaural speech separation. At its core is a content-aware sliding-window attention mechanism that dynamically adapts receptive fields for stronger local interactions, avoiding the rigidity of static convolutions. Unlike time-domain chunk-based methods, TF-MossFormer leverages the 2D spectrogram to model structure along both time and frequency axes. Convolutional gating between attention layers further improves feature selection and information flow. TF-MossFormer achieves SI-SDRi of 22.6, 24.0, and 24.4 dB on WSJ0-2Mix with 5.9M, 16.9M, and 25.4M parameters, respectively, outperforming prior approaches.

cs.SD

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descriptions, and musical attributes. The proposed framework supports three tasks: Lyrics-to-Song Generation, which generates complete songs from text descriptions, lyrics, and musical attributes; Instrumental Music Generation, which creates music without vocals; and Cover Song Generation, which reinterprets existing songs with different styles while preserving their melodic content. Architecturally, our system consists of four main components: a semantic-aware tokenizer, hybird-LM, FullDiT, and a two-level melody module. The tokenizer encodes audio into 8-codebook RVQ tokens for efficient discrete music representation. Based on these tokens, hybird-LM performs hierarchical autoregressive audio-token modeling for full-song generation. To improve audio fidelity, FullDiT performs full-song flow matching in a continuous VAE latent space conditioned on codec tokens, lyrics, and text captions. For cover song generation, the melody module extracts and discretizes melody cues from reference audio to guide generation while preserving the original melodic content. Finally, we investigate DPO, GRPO, and OPD as reward-based post-training strategies for hybird-LM and apply flow-based GRPO to FullDiT to improve musicality and rendering quality. Experimental results on a multilingual automatic benchmark, complemented by the Artificial Analysis Music with Vocals leaderboard, show that the proposed framework achieves competitive performance in the evaluated settings.

cs.SD

Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models

Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attractors -- is a persistent pathology that existing decoding-time heuristics fail to address at its root cause. We present a principled framework that penalises or compensates anomalous confidence arising from collapsed generation patterns, by comparing a token's observed frequency against its corpus prior through an adjacent-conditional probability construction. The resulting self-normalising penalty ratio $R=f(m,n,p)/f(np,n,p)$ requires no ad hoc standardisation and admits a closed-form logit offset with zero approximation error. The correction is isolated from the loss gradient and accumulated into a frozen output-layer bias via exponential moving average, enabling deployment as a repair mechanism for models that have already collapsed without requiring intrusive modifications to standard training pipelines. Experimental validation on a 1.5B-parameter model demonstrates that the frozen-bias mechanism can rescue a model already trapped in a collapsed attractor, reducing 2-gram repetition from 0.073 to near 0 while preserving generation quality.

cs.AI

Early Near-Infrared Excess and Rapid Disk-Corona Evolution in the Tidal Disruption Event 2024aepd

We present multi-wavelength observations of the tidal disruption event (TDE) 2024aepd, spanning primarily the first $\sim$300 days after discovery. The X-ray spectrum is initially dominated by a thermal disk component accompanied by a hard excess. From $\sim$178 days onward, the spectrum becomes power-law dominated and subsequently hardens, indicating the rapid emergence and strengthening of a hot corona. A prominent near-infrared (NIR) excess is detected as early as $\sim40$ days. Its nearly flat power-law spectrum strongly deviates from the Rayleigh-Jeans tail of the UV-optical blackbody. Although a conventional dust-echo origin cannot be completely ruled out, free-free emission from a reprocessing photospheric envelope provides a more plausible explanation. Moreover, the UV-optical-to-NIR break shifts to higher frequencies as the density-profile index remains nearly constant, implying evolving reprocessing conditions within a broadly unchanged density structure. Together with AT2019azh and TDE 2025abcr, TDE 2024aepd is the third TDE reported to exhibit an early-time NIR excess. A larger sample with early-time NIR coverage is needed to determine whether such excesses are common among TDEs.

astro-ph.HE

$J$ and $H$ band sky brightness measurements from polar day to polar night at Dome A, Antarctica

The near-infrared (NIR) sky brightness is a fundamental parameter for evaluating the performance of ground-based infrared observatories. Dome~A on the Antarctic plateau offers exceptional atmospheric conditions, yet its NIR sky background has not been continuously monitored. We present the first continuous $J/H$-band measurements of the sky background at Dome~A from polar day to polar night, and characterize their median levels and temporal variability. The Antarctic Infrared Binocular Telescope (AIRBT), operating in the $J$ and $H$ bands, obtained continuous fixed-pointing observations from February to May 2024, which were used to measure the NIR sky background. The median sky brightness is $5.2/2.9$ and $15.3/13.4~\mathrm{mag~arcsec^{-2}}$ in $J/H$ bands during daytime and nighttime, respectively. The twilight--nighttime boundaries occur at solar elevations of $-9.3^\circ$ in $J$ and $-7.4^\circ$ in $H$. At the same solar elevation, the NIR sky background during the polar night is darker by about $0.1$ and $0.4~\mathrm{mag~arcsec^{-2}}$ in the $J$ and $H$ bands compared with the period of regular day--night alternation. During the polar-night period, the nighttime sky brightness in the $H$ band shows a more evident association with the sunspot number, while the corresponding trend in the $J$ band is weaker. These results reveal systematic differences in sky background between polar and non-polar environments and between polar night and regular day--night cycles. The measured sky brightness may be elevated, as the observations were conducted near solar maximum, highlighting the importance of long-term monitoring across the solar cycle.

astro-ph.IM

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in isolation, lacking a unified benchmark for domain terminology, age variation, dialects, accents, and low-resource languages, particularly across the Middle East and Southeast Asia, representing over one billion under-evaluated speakers. To address this gap, we introduce GigaSpeechBench, a comprehensive multilingual and multidimensional in-the-wild ASR & AST benchmark comprising 680 hours of human-annotated speech. It features five modules: (1) 12 low-resource Middle Eastern and Southeast Asian languages, plus challenging Japanese and Korean; (2) 6 Chinese dialects; (3) 6 English accents; (4) dense terminology across 12 vertical domains for Chinese and English; and (5) older adult and child speech. We further provide human-annotated Chinese and English translations for 11 languages to support AST evaluation. Extensive evaluations of leading foundation models and commercial APIs reveal significant performance degradation in these challenging settings, exposing critical evaluation blind spots.

eess.AS

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis that no current paradigm fully resolves. To address this challenge, we present AudioCALM, a universal audio generation framework that extends autoregressive (AR) next-token prediction from discrete tokens to continuous audio latents: a thin flow-matching head replaces the softmax to predict rectified-flow velocities at each position, and a block-causal AR-Flow attention pattern produces arbitrary-length output. Joint training of multiple audio generation tasks faces an asymmetric text--audio mismatch: speech transcripts align to specific time spans and demand tight, time-aligned attention, whereas sound and music captions describe only overall semantics and rely on diffuse, holistic attention; mixing the two disproportionately degrades sound and music generation. We address this asymmetry at two levels: a data reformulation strategy that unifies all three tasks under a single description-style conditioning interface, and a novel architecture Asymmetric Mixture-of-Modality-Experts (A-MoME), which adds a dedicated residual expert for speech while sound and music share the backbone, incurring no inference overhead on non-speech inputs. Experimental results demonstrate that AudioCALM matches modality-specific state-of-the-art and outperforms prior unified baselines on speech, sound, and music generation benchmarks.

eess.AS

Comparative analysis of missing data imputation methods for CSST survey: Impact on photometric redshift estimation performance

Improving the accuracy of photometric redshifts (photo-$z$) is essential for reliable statistical studies of cosmology and galaxy evolution. However, missing photometric bands are a common observational challenge that can significantly degrade photo-$z$ estimation accuracy. In this work, we present a systematic evaluation of data imputation methods aimed at improving photo-$z$ performance. We benchmark a range of representative machine learning (ML) and deep learning (DL) architectures, identifying k-nearest neighbors (KNN) and the attention-based SAITS model as the leading performers. These models are then applied to China Space Station Survey Telescope (CSST) mock data to assess their performance under realistic observational conditions. Our results show that KNN yields the highest accuracy under idealized missing completely at random (MCAR) conditions with complete training sets, whereas robustness tests reveal that SAITS significantly outperforms KNN when training data is incomplete or when applied to realistic mixed-mechanism scenarios. We find that domain consistency between training and testing missingness patterns is a prerequisite for optimal performance, highlighting the risks of domain shift in supervised regression tasks. Furthermore, our analysis demonstrates that while general imputation models are highly effective for MCAR and missing at random (MAR) data, they are detrimental when applied to missing not at random (MNAR) data arising from flux limits, as statistical models fail to capture the physical information inherent in these non-detections. Consequently, we advocate for more sophisticated architectures capable of disentangling stochastic missingness from physical non-detections to address these distinct mechanisms individually.

astro-ph.GA

Atmospheric turbulence profiling with the Multistar Turbulence Monitor

Accurate characterization of atmospheric optical turbulence is essential for evaluating astronomical sites and optimizing adaptive optics systems. The Multistar Turbulence Monitor (MTM) infers the vertical distribution of the refractive-index structure constant Cn2(z) from differential image motion measured between multiple stellar pairs in short-exposure frames. We present a comprehensive investigation of the MTM method, combining theoretical analysis, instrument-performance assessment, numerical simulations, and on-sky observations obtained at the Daocheng Astronomical Site. Simulations based on a standard HV turbulence model demonstrate that the inversion pipeline robustly recovers both the integrated seeing and the vertical turbulence profile under realistic centroiding noise and varying pixel scales. The Markov Chain Monte Carlo (MCMC) inversion achieves stable results with thirteen discrete height nodes and provides reliable uncertainties. Three nights of MTM measurements at the Daocheng Astronomical Site show that MTM-derived seeing closely tracks simultaneous Differential Image Motion Monitor (DIMM) results, accurately reproducing both short-term fluctuations and nightly averages. These results confirm that MTM provides a simple, portable, and versatile solution for atmospheric turbulence profiling and routine seeing monitoring.

astro-ph.IM

Design, Testing, and Commissioning of the Sun Yat-sen University (SYSU) 80 cm Infrared Telescope

The Sun Yat-sen University (SYSU) 80 cm telescope is a new generation near-infrared (NIR) facility in China dedicated to time-domain astronomy, while also serving as a testbed for emerging NIR cameras. Commissioned in October 2024 at the 4100 m Lenghu site on the Tibetan Plateau in China, the telescope adopts a reflective Cassegrain design with two Nasmyth foci for J and K bands. The J band imaging system, initially equipped with a 640 x 512 off-the-shelf InGaAs camera (INS Mars640) and upgraded in June 2025 to a 1280 x 1024 science-grade, deeply cooled camera (YNAOIR), achieves background-limited performance with a dark current of ~ 14 e-/s/pix and a readout noise of ~ 11 e-. The system reaches a limiting magnitude of J ~ 17 mag (Vega system) in single 20 s exposures and depths of J ~ 19.4 mag with stacked 30 minute exposures. For a variable with J ~ 14 mag during on-sky tests, the system delivers millimagnitude-level photometric precision. Since commissioning, the telescope observed transients such as gamma-ray bursts (GRBs), supernovae and comets, variables including active galactic nuclei (AGNs), high-redshift quasars (z > 6), and brown dwarfs, as well as deep-field imaging reaching J ~ 20.5 mag. This validates the feasibility of using InGaAs cameras for astronomical observations, encouraging other institutions to develop dedicated infrared telescopes or integrate infrared cameras into existing optical telescopes.

astro-ph.IM

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism

Diffusion Transformers (DiTs) are increasingly adopted in scientific computing, yet growing model sizes and resolutions make distributed multi-GPU inference essential. Ulysses sequence parallelism scales DiT inference but introduces frequent all-to-all collectives that dominate latency. Overlapping these with computation is difficult due to tight data dependencies, large message volumes, and asymmetric interconnect bandwidths. We introduce CoCoDiff, a distributed DiT inference engine exploiting two observations: (1) V requires only linear projection while Q/K need additional normalization and RoPE, creating opportunities to overlap V's communication with Q/K computation; (2) adjacent denoising steps produce similar tensors, yielding temporal redundancy. CoCoDiff introduces three mechanisms: Tile-Aware Parallel All-to-all (TAPA) decomposes collectives into topology-aligned phases; V-First scheduling hides V's communication behind Q/K computation; and V-Major selective communication transmits only active projections on slow interconnects. On the Aurora supercomputer with four DiT models across 1-8 nodes (up to 96 Intel GPU tiles), CoCoDiff achieves an average speedup of 3.6x, peaking at 8.4x.

cs.DC

TierBPF: Page Migration Admission Control for Tiered Memory via eBPF

Existing software-based memory tiering systems decide which pages to place on the slower or faster tier. However, they do not take into account two important factors that greatly influence application performance: the size of the migrated pages, and the underlying hardware device and tiering topology. We introduce TierBPF, a software mechanism that can be plugged into existing memory tiering systems to take these factors into account, by making simple binary page admission decisions. TierBPF is implemented as a set of eBPF hooks, which allow users to define their own custom policies. In order to make its decisions, TierBPF utilizes a lightweight tracking mechanism for page profiling which is not dependent on the application's working set size. TierBPF, integrated into three memory tiering systems and evaluated with 17 workloads, achieves geomean throughput gains of up to 17.7% with improvements of up to 75% for individual workloads.

cs.OS

A Robust Geometric Distortion Solution for Main Survey Camera of CSST

The advancement in sensitivity and field of view of next-generation wide-field survey telescopes requires astrometric measurements with high precision, even in the presence of significant geometric distortions. To address this challenge, we develop a Weighted Polynomial Distortion Correction in 2-Phase (WPDC-2P) method. This approach enhances stellar cross-matching, incorporates distance-based weighting into the traditional polynomial fitting, and employs a look-up table to absorb the remaining distortion residuals. Validated on simulated data from the Main Survey Camera of the \emph{Chinese Space Station Survey Telescope} (CSST), incorporating geometric distortions up to approximately $200$ pixels, the method achieves astrometric standard deviation ranging from 0.013 to 0.107 pixels (0.03 pixels for the $g$-1 detector) across all 18 detectors. Under extreme crowding conditions (e.g., globular cluster NGC 2298), the astrometric precision for the $g$-1 detector reaches 0.05-pixel level within the central region ($r_d < 4000$), despite a centroiding precision of $\sim$0.04 pixels. When applied to the Beijing-Arizona Sky Survey data, for which the standard pipeline delivers an astrometric uncertainty of $\sim$20 mas, our method reduces the positional scatter to $ \sigma_{\Delta\alpha}=5.494$ mas (0.01 pixels) and $ \sigma_{\Delta\delta}=9.981$ mas (0.02 pixels) using only a weighted 3rd-order polynomial correction. The method has been integrated into the CSST data processing pipeline and is prepared for further refinement using on-orbit calibration data.

astro-ph.IM

Poisoning the Inner Prediction Logic of Graph Neural Networks for Clean-Label Backdoor Attacks

Graph Neural Networks (GNNs) have achieved remarkable results in various tasks. Recent studies reveal that graph backdoor attacks can poison the GNN model to predict test nodes with triggers attached as the target class. However, apart from injecting triggers to training nodes, these graph backdoor attacks generally require altering the labels of trigger-attached training nodes into the target class, which is impractical in real-world scenarios. In this work, we focus on the clean-label graph backdoor attack, a realistic but understudied topic where training labels are not modifiable. According to our preliminary analysis, existing graph backdoor attacks generally fail under the clean-label setting. Our further analysis identifies that the core failure of existing methods lies in their inability to poison the prediction logic of GNN models, leading to the triggers being deemed unimportant for prediction. Therefore, we study a novel problem of effective clean-label graph backdoor attacks by poisoning the inner prediction logic of GNN models. We propose BA-Logic to solve the problem by coordinating a poisoned node selector and a logic-poisoning trigger generator. Extensive experiments on real-world datasets demonstrate that our method effectively enhances the attack success rate and surpasses state-of-the-art graph backdoor attack competitors under clean-label settings. Our code is available at https://anonymous.4open.science/r/BA-Logic

cs.LG

Beyond Lips: Integrating Gesture and Lip Cues for Robust Audio-visual Speaker Extraction

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. These gestures provide complementary visual cues that can be especially valuable when facial or lip regions are occluded or distant. In this work, we move beyond lip-centric approaches and propose SeLG, a model that integrates both lip and upper-body gesture information for robust speaker extraction. SeLG features a cross-attention-based fusion mechanism that enables each visual modality to query and selectively attend to relevant speech features in the mixture. To improve the alignment of gesture representations with speech dynamics, SeLG also employs a contrastive InfoNCE loss that encourages gesture embeddings to align more closely with corresponding lip embeddings, which are more strongly correlated with speech. Experimental results on the YGD dataset, containing TED talks, demonstrate that the proposed contrastive learning strategy significantly improves gesture-based speaker extraction, and that our proposed SeLG model, by effectively fusing lip and gesture cues with an attention mechanism and InfoNCE loss, achieves superior performance compared to baselines, across both complete and partial (i.e., missing-modality) conditions.

eess.AS

LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial information, including sound source location. This spatial context is crucial for understanding and modeling complex auditory scenes, as it inherently informs sound detection and extraction. In this work, we propose a language-driven universal sound extraction network that isolates text-described sound events from binaural mixtures by effectively leveraging the spatial cues present in binaural signals. Additionally, we jointly predict the direction of arrival (DoA) of the target sound using spatial features from the extraction network. This dual-task approach exploits complementary location information to improve extraction performance while enabling accurate DoA estimation. Experimental results on the in-the-wild AudioCaps dataset show that our proposed LuSeeL model significantly outperforms single-channel and uni-task baselines.

eess.AS

FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning

Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-training alignment via reinforcement learning (RL) effectively aligns generative models with human preferences and downstream metrics in domains such as natural language processing, but its use in speech enhancement remains limited, especially for online RL. Prior work explores offline methods like Direct Preference Optimization (DPO); online methods such as Group Relative Policy Optimization (GRPO) remain largely uninvestigated. In this paper, we present the first successful integration of online GRPO into a flow-matching speech enhancement framework, enabling efficient post-training alignment to perceptual and task-oriented metrics with few update steps. Unlike prior GRPO work on Large Language Models, we adapt the algorithm to the continuous, time-series nature of speech and to the dynamics of flow-matching generative models. We show that optimizing a single reward yields rapid metric gains but often induces reward hacking that degrades audio fidelity despite higher scores. To mitigate this, we propose a multi-metric reward optimization strategy that balances competing objectives, substantially reducing overfitting and improving overall performance. Our experiments validate online GRPO for speech enhancement and provide practical guidance for RL-based post-training of generative audio models.

eess.AS