SearcharxivSearch

arXiv subjects

Xingchen Li

Publications and source records attributed to Xingchen Li.

At least 19 recordsLinked to original sources

A Novel Approach to 3D Dust Mapping of the Central Molecular Zone

The 3D distribution of dust and gas in the Milky Way's Central Molecular Zone (CMZ) is key to understanding gas inflows toward the Galactic Centre (GC), the process of star formation in this extreme environment, and the propagation of energetic cosmic rays originating from Sgr A*. However, while recent efforts have combined datasets in a Bayesian framework to estimate the near/far positions of individual molecular clouds in the CMZ, conflicts between different methodologies still remain and we are still lacking a comprehensive, model-independent map of all of the gas and dust in the CMZ, which is critical to address key science questions. Here we develop a new methodology to infer the 3D dust distribution of the CMZ. The key idea of the method is to use \emph{stellar} proper motions to get probabilistic information about the unknown stellar distances through a model of the distribution of star positions and velocities of the nuclear stellar disc (NSD), co-spatial to the CMZ. Taking \emph{stellar} proper motions and extinctions as input, the latter adopted as a proxy of the dust column density, the method returns the 3D dust distribution. It is non parametric, makes no a-priori assumption on the dust distribution, and is fundamentally distinct and largely independent of all existing methods. We show that the method can robustly and effectively reconstruct the mock 3D CMZ structure by testing it on a range of mock dust distributions, both analytically generated and taken from hydrodynamical simulations. Finally, we discuss the prospects for applying the method to real data.

astro-ph.GA

Scaling Representation Diversity: Modulated Attention and Reconstructive Regularization for Visual Grounding

Referring Expression Comprehension (REC) is commonly studied under dataset-specific fine-tuning, resulting in specialist models with limited cross-dataset generalization. In this work, we revisit REC from the perspective of unified open-vocabulary grounding and identify representation degeneration as a key obstacle to scaling a single generalist model. To preserve representation diversity, we propose a holistic data-model co-design framework. Architecturally, we introduce the Modulated Attention-Contrastive Head (mACH) for efficient token-level vision-language alignment and a text-conditioned JEPA auxiliary stream that provides complementary gradient support to preserve alignment-active representations without inference overhead. On the data side, we introduce Objects365-Caption, enriching Objects365 with context-aware referring expressions for large-scale language supervision. We further provide a theoretical analysis showing that complementary gradient subspaces preserve alignment capacity and thereby scale representation diversity. Extensive experiments demonstrate that our single-checkpoint framework achieves highly competitive performance on standard REC benchmarks while exhibiting strong generalization across heterogeneous grounding datasets without benchmark-specific adaptation.

cs.CV

G-MaP-SE: Guided Speech Enhancement via GMM-Based Prior Matching

Using speaker embeddings as conditioning can strengthen speech enhancement, but most methods either require clean enrollment audio or rely on embeddings extracted from noisy speech, which are fragile under noise and domain shift. We propose G-MaP-SE, a guided enhancement framework that builds a clean-speech embedding prior with a Gaussian Mixture Model (GMM) and refines a noisy conditioning embedding by matching it to this prior. The matched prior embedding is then injected into a time-frequency enhancement backbone via a lightweight gated fusion module. Experiments on VoiceBank+DEMAND and DNS Challenge 2020 datasets show that the proposed prior matching consistently outperforms noisy conditioning and substantially narrows the gap to an oracle clean-conditioning upper bound, while requiring no enrollment audio at inference time. The code, audio samples, and checkpoint are available.

eess.AS

APX-Hardness of Computing Lipschitz Constants for Multi-Parametric Quadratic Programs

Computing the Lipschitz constant of the solution map of a multi-parametric quadratic program is important for the analysis of optimization-based control. This problem is governed by three factors: the parameter dimension, the number of decision variables, and the number of constraints. While empirical evidence has long suggested exponential complexity, a rigorous complexity-theoretic proof has been lacking. In this paper, we fill this gap by proving that this problem is not only NP-hard but also APX-hard. Furthermore, we reveal that: (a) the problem becomes polynomial-time solvable when the number of constraints or decision variables is fixed; and (b) both NP-hardness and APX-hardness persist even in the scalar parameter case. These results confirm that the complexity stems from the number of constraints and variables, rather than the parameter dimension. Numerical experiments further validate these theoretical findings.

eess.SY

Kinematic hints of a nuclear bar in the Milky Way

The Milky Way hosts a flattened nuclear stellar disc (NSD) that dominates the gravitational potential in the inner few hundred parsecs. Whether the NSD is purely axisymmetric or contains a nuclear bar remains an open question. We test for the presence of a nuclear bar using kinematic diagnostics by combining line-of-sight velocities from the KMOS NSD survey with proper motions from VIRAC2 to construct the $ (v_\ell, v_\mathrm{los}) $ velocity ellipse. After applying strict quality cuts to minimise contamination from large-scale bar stars, we measure the vertex deviation $ l_v $ and anisotropy $ \beta $ for several subsamples. For our primary sample ($ |\ell| < 0.9^\circ $, $ -0.4^\circ < b < 0.25^\circ $, $ \mathrm{[Fe/H]} > -0.3 $), we find a significant negative vertex deviation $ l_v = -54.8^{+13.1}_{-14.8}\,^\circ $ with moderate anisotropy $ \beta = 0.16^{+0.08}_{-0.05} $. A subsample restricted to the innermost four fields yields an even stronger signal with $ l_v = -64.3^{+12.1}_{-12.2}\,^\circ $ and $ \beta = 0.38^{+0.12}_{-0.07} $. The direction of maximum velocity dispersion is oriented along Galactic longitude, opposite to that observed in large-scale bar-dominated samples. These signatures are robust against extinction-driven incompleteness, primary-bar contamination, and the choice of metallicity threshold. They are inconsistent with an axisymmetric NSD or one oriented orthogonally to the primary bar, but match expectations for a nuclear bar oriented at $ \alpha \approx 60^\circ $-$75^\circ$ to the Sun-Galactic-Centre line with its near side pointing toward positive Galactic longitude. While definitive confirmation awaits larger and more precise samples from upcoming surveys, our results provide the first kinematic indication of a possible nuclear bar in the Milky Way.

astro-ph.GA

Simulations of gas inflow in the Milky Way I. Stellar-Feedback-Regulated Transport from the Central Molecular Zone to the Circumnuclear disk

We perform hydrodynamical simulations with radially varying resolution to study the effects of stellar feedback on the radial inflow of gas from the Central Molecular Zone (CMZ, $R\sim200$ pc) to the Circumnuclear Disk (CND, $R\sim5$ pc) of the Milky Way. The simulations include a realistic Milky Way barred gravitational potential, a cooling function coupled to a non-equilibrium chemical network, gas self-gravity, star formation, supernova feedback, and radiation feedback from massive stars computed via on-the-fly radiative transfer. Our main findings are as follows: 1) Stellar feedback drives a radial inflow that decreases monotonically with decreasing Galactocentric radius. The time-averaged inflow rate in our fiducial SNRad simulation, which includes both supernova and radiation feedback, declines from $\langle \dot{M} \rangle\sim5\times10^{-3}$ Msun/yr at $R\sim100$ pc, to $\langle\dot{M}\rangle\sim10^{-4}$ Msun/yr at $R\sim10$ pc, to $\langle\dot{M}\rangle\sim10^{-6}$ Msun/yr at $R\sim1$ pc. 2) The total inflow rate can be broken down into two components driven by two distinct mechanisms. First, feedback-driven turbulence redistributes the angular momentum of gas clouds, producing a smooth (secular) transport of mass inward, similar to a Shakura-Sunyaev viscous accretion disk. This component contributes inflow rates that vary from $\dot{M}\sim5\times10^{-4}$ Msun/yr at $R\sim100$ pc to $\dot{M}\sim10^{-7}$ Msun/yr at $R\sim1$ pc. Second, episodic inflow events can transiently increase the inflow rate by several orders of magnitude, reaching $\dot{M}\sim10^{-3}$ Msun/yr over timescales of $\Delta t\sim3$-$5$ Myr at $R=10$ pc. 3) The stellar feedback model significantly affects the episodic inflow but has little impact on the smooth component. Simulations including radiation feedback produce substantially more episodic events than those with supernova feedback alone.

astro-ph.GA

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafted heuristics, and now machine learning has emerged as a promising paradigm to automate this process. However, existing learning-based methods are often hindered by their dependence on expensive expert demonstrations and the gap between training objectives and the solver's end-to-end performance. In this work, we propose LLM4Branch, a novel framework that leverages Large Language Models (LLMs) to automate the discovery of efficient branching policies. Specifically, the discovered policy is an executable program with a program skeleton generated by the LLM and a parameter vector, which is optimized via a zeroth-order method over a few instances with their end-to-end performance feedback. Extensive experiments on standard MILP benchmarks demonstrate that LLM4Branch establishes a new state-of-the-art among CPU-based methods and achieves performance competitive with advanced GPU-based models. Codes are available at https://github.com/hzn18/LLM4Branch.

cs.AI

EvoTSE: Evolving Enrollment for Target Speaker Extraction

Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker confusion, where models mistakenly extract the interfering speaker. Furthermore, conventional TSE relies on static inference pipeline, where performance is limited by the quality of the fixed enrollment. To overcome these limitations, we propose EvoTSE, an evolving TSE framework in which the enrollment is continuously updated through reliability-filtered retrieval over high-confidence historical estimates. This mechanism reduces speaker confusion and relaxes the quality requirements for pre-recorded enrollment without relying on additional annotated data. Experiments across multiple benchmarks demonstrate that EvoTSE achieves consistent improvements, especially when evaluated on out-of-domain (OOD) scenarios. Our code and checkpoints are available.

eess.AS

Kinematic diagnostics for non-axisymmetry in the Milky Way's nuclear stellar disc

There is now strong evidence that the Milky Way (MW) hosts a nuclear stellar disc (NSD). However, whether the NSD is purely axisymmetric or contains a nuclear bar remains unresolved. Since approximately $50\%$ of barred galaxies with MW-like mass in the local Universe host a nuclear bar, investigating whether the MW hosts one is of interest. We conduct a systematic analysis to identify robust kinematic diagnostics capable of determining whether the MW hosts a nuclear bar. Using N-body simulations, we explore the kinematic signatures indicative of a nuclear bar. Using the phase-space coordinates longitude $(\ell)$, latitude $(b)$, proper motions ($\mu_\ell$ and $\mu_{\rm b})$ and line-of-sight velocity $(v_{\rm los})$, we test various diagnostics assuming different nuclear bar orientations. We also evaluate how sample size, dust extinction and bar amplitude influence the efficacy of the diagnostics. We identify two independent kinematic diagnostics capable of revealing a nuclear bar in the MW: (1) the vertex deviation, $l_{\rm v}$, of the ($v_{\ell}-v_{\rm los}$) velocity ellipse; and (2) The asymmetry in the $\mu_{\ell}$ vs $\ell$ distribution. While both are impacted by the sample size and extinction, the vertex deviation proves more robust, especially when combining stars from multiple observational fields. We also assess the correlation between the line-of-sight velocity and the $h_3$ Gauss-Hermite moment ("skewness") of the line-of-sight velocity but find no clear distinction between an NSD and a nuclear bar based on this metric. Our results suggest that data from the current KMOS survey may allow a marginal detection of a nuclear bar using the vertex deviation method. A companion paper provides further validation and detailed analysis of this approach. Nonetheless, future surveys will provide the high quality data necessary to fully exploit the diagnostics outlined in this study.

astro-ph.GA

DiT-IC: Aligned Diffusion Transformer for Efficient Image Compression

Diffusion-based image compression has recently shown outstanding perceptual fidelity, yet its practicality is hindered by prohibitive sampling overhead and high memory usage. Most existing diffusion codecs employ U-Net architectures, where hierarchical downsampling forces diffusion to operate in shallow latent spaces (typically with only 8x spatial downscaling), resulting in excessive computation. In contrast, conventional VAE-based codecs work in much deeper latent domains (16x - 64x downscaled), motivating a key question: Can diffusion operate effectively in such compact latent spaces without compromising reconstruction quality? To address this, we introduce DiT-IC, an Aligned Diffusion Transformer for Image Compression, which replaces the U-Net with a Diffusion Transformer capable of performing diffusion in latent space entirely at 32x downscaled resolution. DiT-IC adapts a pretrained text-to-image multi-step DiT into a single-step reconstruction model through three key alignment mechanisms: (1) a variance-guided reconstruction flow that adapts denoising strength to latent uncertainty for efficient reconstruction; (2) a self-distillation alignment that enforces consistency with encoder-defined latent geometry to enable one-step diffusion; and (3) a latent-conditioned guidance that replaces text prompts with semantically aligned latent conditions, enabling text-free inference. With these designs, DiT-IC achieves state-of-the-art perceptual quality while offering up to 30x faster decoding and drastically lower memory usage than existing diffusion-based codecs. Remarkably, it can reconstruct 2048x2048 images on a 16 GB laptop GPU.

eess.IV

ReLU Networks for Model Predictive Control: Network Complexity and Performance Guarantees

Recent years have witnessed a resurgence in using ReLU neural networks (NNs) to represent model predictive control (MPC) policies. However, determining the required network complexity to ensure closed-loop performance remains a fundamental open problem. This involves a critical precision-complexity trade-off: undersized networks may fail to capture the MPC policy, while oversized ones may outweigh the benefits of ReLU network approximation. In this work, we propose a projection-based method to enforce hard constraints and establish a state-dependent Lipschitz continuity property for the optimal MPC cost function, which enables sharp convergence analysis of the closed-loop system. For the first time, we derive explicit bounds on ReLU network width and depth for approximating MPC policies with guaranteed closed-loop performance. To further reduce network complexity and enhance closed-loop performance, we propose a non-uniform error framework with a state-aware scaling function to adaptively adjust both the input and output of the ReLU network. Our contributions provide a foundational step toward certifiable ReLU NN-based MPC.

eess.SY

YODA: Yet Another One-step Diffusion-based Video Compressor

While one-step diffusion models have recently excelled in perceptual image compression, their application to video remains limited. Prior efforts typically rely on pretrained 2D autoencoders that generate per-frame latent representations independently, thereby neglecting temporal dependencies. We present YODA--Yet Another One-step Diffusion-based Video Compressor--which embeds multiscale features from temporal references for both latent generation and latent coding to better exploit spatial-temporal correlations for more compact representation, and employs a linear Diffusion Transformer (DiT) for efficient one-step denoising. YODA achieves state-of-the-art perceptual performance, consistently outperforming traditional and deep-learning baselines on LPIPS, DISTS, FID, and KID. Source code will be publicly available at https://github.com/NJUVISION/YODA.

eess.IV

LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation

In-car multi-zone speech separation, which captures voices from different speech zones, plays a crucial role in human-vehicle interaction. Although previous SpatialNet has achieved notable results, its high computational cost still hinders real-time applications in vehicles. To this end, this paper proposes LSZone, a lightweight spatial information modeling architecture for real-time in-car multi-zone speech separation. We design a spatial information extraction-compression (SpaIEC) module that combines Mel spectrogram and Interaural Phase Difference (IPD) to reduce computational burden while maintaining performance. Additionally, to efficiently model spatial information, we introduce an extremely lightweight Conv-GRU crossband-narrowband processing (CNP) module. Experimental results demonstrate that LSZone, with a complexity of 0.56G MACs and a real-time factor (RTF) of 0.37, delivers impressive performance in complex noise and multi-speaker scenarios.

cs.SD

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement

Generative Universal Speech Enhancement (USE) methods aim to leverage generative models to improve speech quality under various types of distortions. However, existing generative speech enhancement methods often suffer from semantic inconsistency in the generated outputs. Therefore, we propose SenSE, a novel two-stage generative universal speech enhancement framework, by modeling semantic priors with a language model, the flow matching-based speech enhancement process is guided to generate semantically faithful speech, thereby effectively improving context fidelity. In addition, we introduce a dual-path masked conditioning training strategy that enables flow matching-based enhancement to flexibly integrate multi-source conditioning signals from degraded speech, semantic tokens, and reference speech, thereby improving model flexibility and adaptability. Experimental results demonstrate that SenSE achieves state-of-the-art performance among generative speech enhancement models and exhibits a high performance ceiling, particularly under challenging distortion conditions. Codes and demos are available at https://github.com/ASLP-lab/SenSE.

eess.AS

MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow

Speech enhancement (SE) recovers clean speech from noisy signals and is vital for applications such as telecommunications and automatic speech recognition (ASR). While generative approaches achieve strong perceptual quality, they often rely on multi-step sampling (diffusion/flow-matching) or large language models, limiting real-time deployment. To mitigate these constraints, we present MeanFlowSE, a one-step generative SE framework. It adopts MeanFlow to predict an average-velocity field for one-step latent refinement and conditions the model on self-supervised learning (SSL) representations rather than VAE latents. This design accelerates inference and provides robust acoustic-semantic guidance during training. In the Interspeech 2020 DNS Challenge blind test set and simulated test set, MeanFlowSE attains state-of-the-art (SOTA) level perceptual quality and competitive intelligibility while significantly lowering both real-time factor (RTF) and model size compared with recent generative competitors, making it suitable for practical use. The code will be released upon publication at https://github.com/Hello3orld/MeanFlowSE.

cs.SD

UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling

Generative modeling has recently achieved remarkable success across image, video, and audio domains, demonstrating powerful capabilities for unified representation learning. Yet speech front-end tasks such as speech enhancement (SE), target speaker extraction (TSE), acoustic echo cancellation (AEC), and language-queried source separation (LASS) remain largely tackled by disparate, task-specific solutions. This fragmentation leads to redundant engineering effort, inconsistent performance, and limited extensibility. To address this gap, we introduce UniFlow, a unified framework that employs continuous generative modeling to tackle diverse speech front-end tasks in a shared latent space. Specifically, UniFlow utilizes a waveform variational autoencoder (VAE) to learn a compact latent representation of raw audio, coupled with a Diffusion Transformer (DiT) that predicts latent updates. To differentiate the speech processing task during the training, learnable condition embeddings indexed by a task ID are employed to enable maximal parameter sharing while preserving task-specific adaptability. To balance model performance and computational efficiency, we investigate and compare three generative objectives: denoising diffusion, flow matching, and mean flow within the latent domain. We validate UniFlow on multiple public benchmarks, demonstrating consistent gains over state-of-the-art baselines. UniFlow's unified latent formulation and conditional design make it readily extensible to new tasks, providing an integrated foundation for building and scaling generative speech processing pipelines. To foster future research, we will open-source our codebase.

eess.AS

EchoFree: Towards Ultra Lightweight and Efficient Neural Acoustic Echo Cancellation

In recent years, neural networks (NNs) have been widely applied in acoustic echo cancellation (AEC). However, existing approaches struggle to meet real-world low-latency and computational requirements while maintaining performance. To address this challenge, we propose EchoFree, an ultra lightweight neural AEC framework that combines linear filtering with a neural post filter. Specifically, we design a neural post-filter operating on Bark-scale spectral features. Furthermore, we introduce a two-stage optimization strategy utilizing self-supervised learning (SSL) models to improve model performance. We evaluate our method on the blind test set of the ICASSP 2023 AEC Challenge. The results demonstrate that our model, with only 278K parameters and 30 MMACs computational complexity, outperforms existing low-complexity AEC models and achieves performance comparable to that of state-of-the-art lightweight model DeepVQE-S. The audio examples are available.

eess.AS

4H-SiC PIN detector for alpha particles from room temperature to 90 {\deg}C

In the field of high-energy particle detection, detectors operating in high-radiation environments primarily face high costs associated with power consumption and cooling systems. Therefore, the development of particle detectors capable of stable operation at room temperature or even elevated temperatures is of great significance. Silicon carbide (SiC) exhibits significant potential for particle detector applications due to its exceptional carrier mobility, radiation hardness, and thermal stability. Over the past decade, significant breakthroughs in silicon carbide epitaxial growth technology and device processing techniques have enabled the development of SiC-based particle detectors, providing a new technological pathway for particle detection in high-temperature environments. In this work, we fabricate a 4H-SiC PIN detector, named SIlicon CARbide (SICAR) and characterize its leakage current, capacitance, and charge collection across varying temperatures. The results indicate that the detector maintains a very low leakage current (< 10 nA) at 90 C, with no degradation in depletion capacitance or charge collection performance. Additionally, it achieves a fast rise time of 333 ps at 90 C, confirming its potential for high-temperature radiation detection applications.

physics.ins-det