SearcharxivSearch

arXiv subjects

Xian Shi

Publications and source records attributed to Xian Shi.

At least 19 recordsLinked to original sources

Error Exponents of Probabilistic Quantum Resource Distillation

In this manuscript, we establish a unified framework for analyzing the conditional error exponents of probabilistic resource distillation under approximately resource-nongenerating instruments. For generic quantum resource theories satisfying suitable structural conditions, we derive general one-shot bounds on the conditional distillation error exponents. By relating probabilistic distillation to postselected composite quantum hypothesis testing, we further obtain analytical characterizations of the conditional error exponents for entanglement, coherence, and magic distillation at finite blocklength, and derive their asymptotic expressions in the zero-rate regime. For several representative families of states in entanglement and magic resource theories, these characterizations reduce to explicit closed-form formulas. Comparing them with the corresponding deterministic distillation exponents, we identify regimes in which postselection yields a strict improvement in the exponential decay rate of the conditional error. Our results reveal an operational advantage of postselection in quantum resource distillation and establish postselected composite hypothesis testing as a general tool for characterizing probabilistic resource-processing tasks.

quant-ph

Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods

This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.

eess.AS

DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios

Multi-speaker automatic speech recognition (MSASR) aims to jointly predict content transcriptions, speaker identities, and timestamps, thereby addressing the key question of "who spoke what and when" and holds substantial practical value in real-world multi-speaker scenarios. However, MSASR still encounters considerable challenges in the presence of fast turn transitions, overlapping speech, and complex, diverse multi-speaker scenarios. In this work, we propose DiaScriber, an end-to-end multi-speaker diarization and transcription model built on a speech large language model. We first construct diverse data pipelines to cover a wide variety of multi-speaker scenarios and their complexities, including validation and refinement, turn-transition and overlapping-speech simulation, and multimodal annotation. Furthermore, DiaScriber is developed based on the pretrained version of Qwen3.5-Omni through a three-stage training strategy involving continual pretraining, supervised fine-tuning, and reinforcement learning. Experiments show that DiaScriber achieves superior performance over comparison methods across extensive multi-speaker scenario test sets and demonstrates outstanding generalization ability in unseen multi-speaker scenarios.

eess.AS

Quantum Probabilistic Local Differential Privacy: Structural Properties and Sample Complexity Bounds

Differential privacy provides a rigorous framework for quantifying privacy leakage in data analysis, while its quantum extensions have become increasingly relevant with the development of quantum computing and quantum machine learning. In this work, we introduce and study quantum probabilistic local differential privacy, a relaxation of quantum local differential privacy in which the privacy constraint is allowed to fail on a spectral violation event with low probability. This quantity can be interpreted as the probability under the quantum superoperation of a quantum privacy-loss violation, and is closely related to the acceptance probability of the quantum Neyman-Pearson test at a small threshold. We investigate the basic structural properties of this privacy notion and clarify its relationship with existing forms of quantum differential privacy. We show the properties of quantum probabilistic local differential privacy under tensor-product composition and unitary post-processing, while it is in general neither convex nor closed under post-processing by arbitrary quantum channels. We further characterize when depolarizing noise satisfies quantum probabilistic local differential privacy under several representative scenarios. Finally, we connect quantum probabilistic privacy constraints with statistical inference by deriving a lower bound on probabilistically privatized contraction coefficients in terms of the hockey-stick divergence. As an application, we obtain sample complexity bounds of probabilistically privated asymmetric and symmetric quantum hypothesis testing. These results provide a systematic foundation for studying probabilistic privacy guarantees in quantum information processing and their operational consequences for private quantum statistical inference.

quant-ph

Shared Keyboard: An improved Bayesian design for phase I clinical trials via Beta kernel process

Model-assisted interval designs such as the Keyboard design are transparent and easy to implement in phase I oncology trials. However, interim decisions based solely on data from the current dose may overlook informative signals from neighbouring doses, leading to unnecessary escalation or de-escalation. We propose the shared Keyboard design, a Bayesian model-assisted design that replaces the independent beta--binomial updating scheme at each dose with a posterior induced by a Beta kernel process using kernel-weighted pseudo-counts. The design preserves the decision structure of the Keyboard design while enabling controlled borrowing across nearby doses. To prioritise overdose control, we propose an asymmetric kernel that assigns greater weight to toxicities observed at higher doses during escalation. We further extend the proposed design to accommodate adaptive dose insertion when the initial dose grid is inadequate and time-to-event outcomes when late-onset toxicities are present. Extensive simulation studies demonstrate substantial improvements in both accuracy and safety for identifying the maximum tolerated dose. In settings involving dose insertion, the proposed design identifies inserted target doses more effectively than adaptive dose modification while maintaining a comparable modification rate.

stat.AP

Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model

Achieving natural full-duplex interaction in spoken dialogue systems (SDS) remains a challenge due to the difficulty of accurately detecting user interruptions. Current solutions are polarized between "trigger-happy" VAD-based methods that misinterpret backchannels and robust end-to-end models that exhibit unacceptable response delays. Moreover, the absence of real-world benchmarks and holistic metrics hinders progress in the field. This paper presents a comprehensive frame-work to overcome these limitations. We first introduce SID-Bench, the first benchmark for semantic-aware interruption detection built entirely from real-world human dialogues. To provide a rigorous assessment of the responsiveness-robustness trade-off, we propose the Average Penalty Time (APT) metric, which assigns a temporal cost to both false alarms and late responses. Building on this framework, we design an LLM-based detection model optimized through a novel training paradigm to capture subtle semantic cues of intent. Experimental results show that our model significantly outperforms mainstream baselines, achieving a nearly threefold reduction in APT. By successfully resolving the long-standing tension between speed and stability, our work establishes a new state-of-the-art for intelligent interruption handling in SDS. To facilitate future research, SID-Bench and the associated code are available at: https://github.com/xkx-hub/SID-bench.

cs.SD

Pre-perihelion Volatile Evolution of Interstellar Comet 3I/ATLAS Indicating Significant Contribution from Extended Source in the Coma

Interstellar comets provide rare opportunities for probing the diversity of refractory and volatile inventory around other stars. As the second ever interstellar comet, and the third interstellar object, 3I/ATLAS has been the focus of telescopic observations since its discovery in July 2025. Following the previous observations at multi-wavelengths, we present further radio observations of the 1665/1667 MHz ground-state OH lines and millimeter observations of the CO($J$=1-0) transition at 115.271 GHz that trace the coma $\rm H_2O$ and CO abundances, respectively. We derived OH production rates of $(1.32\pm0.47)\times10^{28}\ \rm s^{-1}$ at 2.27 au and $(1.89\pm0.37)\times10^{28}\ \rm s^{-1}$ at 1.96 au as well as an average CO production rate of $\rm (5.75\pm1.91) \times 10^{27}\, s^{-1}$ between 2.33 and 1.75 au, inferring a CO/$\rm H_2O$ ratio of ($28\pm11\%$). With the mean HCN production rate of $2.5\times 10^{25}\ \rm s^{-1}$ at 2.1 au reported by \citet{2025arXiv251120845R} and \citet{2025arXiv251002817C}, we infer a CO/HCN ratio of ($230\pm76$). By synthesizing water production rates measured with instruments of different apertures, we found that the sublimation from extended source in the coma contributes significantly to 3I's pre-perihelion water measurements, accounting for up to 80\% from 3 au to 2 au.

astro-ph.EP

Qwen3-ASR Technical Report

In this report, we introduce Qwen3-ASR family, which includes two powerful all-in-one speech recognition models and a novel non-autoregressive speech forced alignment model. Qwen3-ASR-1.7B and Qwen3-ASR-0.6B are ASR models that support language identification and ASR for 52 languages and dialects. Both of them leverage large-scale speech training data and the strong audio understanding ability of their foundation model Qwen3-Omni. We conduct comprehensive internal evaluation besides the open-sourced benchmarks as ASR models might differ little on open-sourced benchmark scores but exhibit significant quality differences in real-world scenarios. The experiments reveal that the 1.7B version achieves SOTA performance among open-sourced ASR models and is competitive with the strongest proprietary APIs while the 0.6B version offers the best accuracy-efficiency trade-off. Qwen3-ASR-0.6B can achieve an average TTFT as low as 92ms and transcribe 2000 seconds speech in 1 second at a concurrency of 128. Qwen3-ForcedAligner-0.6B is an LLM based NAR timestamp predictor that is able to align text-speech pairs in 11 languages. Timestamp accuracy experiments show that the proposed model outperforms the three strongest force alignment models and takes more advantages in efficiency and versatility. To further accelerate the community research of ASR and audio understanding, we release these models under the Apache 2.0 license.

cs.CL

LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech

Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The multilingual speech understanding and long-sequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings. However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference. To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript. Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots. During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions. Dynamic slot insertion enables FA at arbitrary positions. Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed. Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69%~78% relative reduction in accumulated averaging shift compared with prior methods. Checkpoint and inference code are available at https://github.com/QwenLM/Qwen3-ASR.

cs.SD

Probabilistic Entanglement Distillation: Error Exponents via Postselected Quantum Hypothesis Testing Against Separable States

Entanglement distillation is a fundamental task in quantum entanglement theory. While recent progress has clarified limitations of probabilistic transformations in general resource theories, an analytic formula for the error exponent of probabilistic entanglement distillation under approximately (dually) nonentangling operations has remained unavailable. This work considers the error exponents of probabilistic entanglement distillation under the finite block length scenario and zero rate scenario when the operational model is $\delta$-approximately nonentangling or $\delta$-approximately dually nonentangling quantum instruments. Building on the framework of postselected quantum hypothesis testing, we establish a direct connection between probabilistic distillation and postselected hypothesis testing against the set of separable states. In particular, we derive an analytical characterization of the distillation error exponent under $\delta$-(dually) approximately nonentangling quantum instruments.

quant-ph

Antarctic TianMu Staring Observation Project I: Overview and Implementation of the Prototype Telescope

Wide-field rapid sky surveys serve as critical observational methods for time-domain astronomical research. The Antarctic region, with several months of continuous dark nights annually, is an ideal site for time-domain astronomical observations. The Antarctic TianMu Staring Observation Project aims to deploy a fleet of small telescopes, adopting an array observation model to conduct time-domain optical observations in Antarctica, featuring wide-sky coverage, high-cadence sampling, long-period staring, and simultaneous multi-band measurements. Considering the severe challenges optical telescopes face in Antarctica, including extremely low temperatures, unattended operation, and limited power supply and network transmission, we have designed and developed the Antarctic TianMu prototype telescope based on drift-scan charge-coupled device technology. In October 2022, our prototype (with an aperture of 18 cm), named AT-Proto was transported to Zhongshan Station in Antarctica aboard China's 39th Antarctic Research Expedition. It has since operated stably and reliably in the frigid environment for over two years, demonstrating the significant advantages of this technology in polar astronomical observations. The experimental observation results of AT-Proto provide a solid foundation for the subsequent construction of a time-domain astronomy observation array in Antarctica.

astro-ph.IM

Qwen3-Omni Technical Report

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts. Qwen3-Omni matches the performance of same-sized single-modal models within the Qwen series and excels particularly on audio tasks. Across 36 audio and audio-visual benchmarks, Qwen3-Omni achieves open-source SOTA on 32 benchmarks and overall SOTA on 22, outperforming strong closed-source models such as Gemini-2.5-Pro, Seed-ASR, and GPT-4o-Transcribe. Qwen3-Omni adopts a Thinker-Talker MoE architecture that unifies perception and generation across text, images, audio, and video, yielding fluent text and natural real-time speech. It supports text interaction in 119 languages, speech understanding in 19 languages, and speech generation in 10 languages. To reduce first-packet latency in streaming synthesis, Talker autoregressively predicts discrete speech codecs using a multi-codebook scheme. Leveraging the representational capacity of these codebooks, we replace computationally intensive block-wise diffusion with a lightweight causal ConvNet, enabling streaming from the first codec frame. In cold-start settings, Qwen3-Omni achieves a theoretical end-to-end first-packet latency of 234 ms. To further strengthen multimodal reasoning, we introduce a Thinking model that explicitly reasons over inputs from any modality. Since the research community currently lacks a general-purpose audio captioning model, we fine-tuned Qwen3-Omni-30B-A3B to obtain Qwen3-Omni-30B-A3B-Captioner, which produces detailed, low-hallucination captions for arbitrary audio inputs. Qwen3-Omni-30B-A3B, Qwen3-Omni-30B-A3B-Thinking, and Qwen3-Omni-30B-A3B-Captioner are publicly released under the Apache 2.0 license.

cs.CL

Erasing, Converting, and Communicating: The Power of Resource-Nongenerating Operations

We investigate resource nongenerating operations in both static and dynamical quantum resource theories. For the static scenarios, we derive a sufficient condition for state transformations under resource nongenerating operations. Then we construct a dynamical resource theory where resource nongenerating operations constitute the set of free operations, and we propose an axiomatic approach to quantify the dynamical resource. We further analyze the erasure of the dynamical resources. As applications, we establish bounds on the rate of state conversions under resource nongenerating operations in a generic convex resource theory and obtain capacity bounds for classical communication tasks assisted by dynamical coherence. Our results clarify the key roles of resource nongenerating operations in quantum information processing tasks.

quant-ph

Dynamically New Comet C/2025 D1 (Groeller) with Record Perihelion Distance

We studied C/2025 D1 (Groeller), a long-period comet with an unprecedented perihelion distance of 14.1 au, using archival observations. The data reveals that it had been active at inbound heliocentric distances $r_{\rm H} \gtrsim 20$ au. Initially, the comet intrinsically brightened at $r_{\rm H} \gtrsim 16$ au, with brightening parameters comparable to those of other long-period comets. However, observations after late 2023 showed a gradual decay, despite the inbound trajectory of the comet. To our knowledge, such behaviours have not been observed for other long-period comets at similar heliocentric distances. We speculate that this might be linked to the onset of CO$_{2}$ sublimation and/or crystallisation processes. Alternatively, the activity source might have been exhausted. The surface brightness profile of the coma indicates a steady-state mass loss, implying supervolatile sublimation as the primary driver of the observed activity. Despite changes in the orbital plane angle, the circularly symmetric coma persisted throughout the observed period, indicative of the dominance of large grains in the coma. Assuming the activity trend is independent of bandpass, we found that comet was redder than many other solar system comets. Our model-dependent constraint estimates the nucleus radius to be $\gtrsim\!0.4$ km. We performed astrometric measurements, refined the orbital solution, and derived the original and future orbits of the comet. Our N-body integration, accounting for the Galactic tide, strongly favours that the comet is dynamically new, with its previous perihelion at $\gtrsim\!60$ au from the Sun $\gtrsim\!6$ Myr ago. It is highly likely that the comet will be lost from our solar system after the current apparition.

astro-ph.EP

Pre-perihelion radio observations of comet 12P/Pons-Brooks with Tianma radio telescope

{The multiple outburst events of comet 12P/Pons-Brooks during its 2024 apparition offer a unique window into highly-active volatile releasing processes not observable during quiescent periods. We performed radio observations of comet 12P/Pons-Brooks with the Tianma-65m radio telescope, targeting the OH and NH$_3$ inversion lines at 18-cm and 1.3-cm, respectively. By monitoring 12P at different heliocentric distances on its inbound journey, we aim to provide insights into the comet's volatile composition and outburst behavior. Four observations were carried out between December 2023 and March 2024 when the comet was approaching the Sun from 2.22 AU to 1.18 AU. We conducted 18-cm OH lines observations on 4 single days using the cryogenically cooled receiver system of the telescope to derive $\rm H_{2}O$ production rate. During 12P's outburst on December 14, we also conducted observations targeting the $\rm NH_{3}$ emission. OH 18-cm lines were clearly detected with a signal-to-noise ratio of $\sim$4$\sigma$ (peak intensity). A tentative detection of $\rm NH_{3}$ was made at the $\sim$$3\sigma$ level during the outburst phase, but the detection needs to be further verified. Our observations provide information on the outgassing behavior of 12P/Pons-Brooks during its 2024 apparition. The water production rate of 12P, derived from the 18-cm OH lines is consistent with measurements obtained in other works. The possible detection of $\rm NH_{3}$ during an outburst suggests possible connections between subsurface volatile reservoir and the outburst mechanism. These results could further our understanding of the composition and activity of Halley-type comets.

astro-ph.EP

Schmidt-number robustness as a unified quantifier of high dimensional entanglement in Buscemi nonlocality

High-dimensional entanglement, captured by the Schmidt number, underpins advantages in quantum information tasks, yet a unified resource-theoretic description across different Buscemi-type operational objects has been missing. Here we develop a convex framework that treats bipartite states, distributed measurements, and teleportation instruments generated from shared entanglement on equal footing. For a fixed Schmidt-number threshold k, we introduce robustness-based monotones for each class of objects and prove a quantitative collapse: the Schmidt-number robustness of a bipartite state coincides with the maximal robustness achievable by any distributed measurement or teleportation instrument derived from that state. Consequently, within Buscemi-type operational frameworks, these objects do not carry independent high-dimensional resources but are governed by a single robustness-based monotone. We further provide a direct operational interpretation by relating this unique quantifier to the optimal advantage in entanglement-assisted state discrimination games. Our results complete a unified resource-theoretic characterization of high-dimensional entanglement across states, measurements, and quantum devices.

quant-ph

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.

cs.SD

ThermoONet -- a deep learning-based small body thermophysical network: applications to modelling water activity of comets

Cometary activity is a compelling subject of study, with thermophysical models playing a pivotal role in its understanding. However, traditional numerical solutions for small body thermophysical models are computationally intensive, posing challenges for investigations requiring high-resolution or repetitive modeling. To address this limitation, we employed a machine learning approach to develop ThermoONet - a neural network designed to predict the temperature and water ice sublimation flux of comets. Performance evaluations indicate that ThermoONet achieves a low average error in subsurface temperature of approximately 2% relative to the numerical simulation, while reducing computational time by nearly six orders of magnitude. We applied ThermoONet to model the water activity of comets 67P/Churyumov-Gerasimenko and 21P/Giacobini-Zinner. By successfully fitting the water production rate curves of these comets, as obtained by the Rosetta mission and the SOHO telescope, respectively, we demonstrate the network's effectiveness and efficiency. Furthermore, when combined with a global optimization algorithm, ThermoONet proves capable of retrieving the physical properties of target bodies.

astro-ph.EP