SearcharxivSearch

arXiv subjects

Jing Lu

Publications and source records attributed to Jing Lu.

At least 19 recordsLinked to original sources

The host galaxies of 2003fg-like type Ia supernovae

SN~2003fg-like events are a peculiar type Ia supernova (SN~Ia) subtype characterized by broader light curves, higher near-infrared luminosities, and stronger carbon absorptions at early times. Here we present observations of the largest compilation of 2003fg-like SN Ia host galaxies to date, obtained with Integral Field Spectroscopy (IFS). For 20 objects, we study both the global host-galaxy properties and, for the first time for a sizeable sample, the local environment at the SN position. Globally, 2003fg-like SNe~Ia occur in galaxies with lower stellar mass, lower oxygen abundance, and marginally higher specific star-formation rate (sSFR) than those of normal SNe~Ia, although their hosts are not as extreme in such properties as superluminous SNe, nor representative of metal-poor dwarf-galaxy samples. Locally, the SN positions show lower star-formation-rate, stellar-mass surface densities, and lower sSFR, than normal SN~Ia environments, consistent with a significant preference for the outskirts of their hosts, while their stellar age indicators are typical. The most distinctive local property is metallicity, with 2003fg-like SNe~Ia occupying the most metal-poor environments among SNe~Ia. We also find a tentative positive correlation between the light-curve width and the oxygen abundance for 2003fg-like events. Our results imply that 2003fg-like SNe~Ia arise from the merger of two white dwarfs (WDs) or the core-degenerate scenario, but disfavor the single, rapidly rotating super-M$_{ch}$ C-O WD progenitor, as this channel requires a young stellar population that we do not observe at the SN positions. The preference of 2003fg-like SNe~Ia for low-metallicity environments suggests that they may have been more common in the early Universe. (abridged)

astro-ph.CO

Feedback-Guided DNN-Based Controller Fusion for Robust Fixed-Parameter Active Noise Control

In active noise control (ANC) systems, adaptive approaches may suffer from instability or divergence, limiting their practical deployment. Consequently, fixed-parameter controllers are widely adopted, but their performance degrades under varying noise characteristics and acoustic path conditions. This paper proposes a feedback-guided DNN-based controller fusion framework for robust fixed-parameter ANC. The proposed method combines a causal WaveNet controller with a feedback-guided mixture-of-experts (MoE) module, where a gating network estimates the weights of multiple pre-trained FIR experts according to the current acoustic condition. The proposed approach improves robustness to varying acoustic conditions without online parameter updating. Furthermore, the model is fully causal and supports sample-wise streaming inference, with computational costs evenly distributed across sampling points to reduce peak computational load. Experimental results on headphone ANC demonstrate substantial low-frequency noise reduction with negligible noise amplification over 1-8 kHz.

eess.SY

CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement

Ultra-lightweight models are essential for the deployment of deep learning-based speech enhancement algorithms on edge devices. Although recent approaches have achieved a certain balance between computational complexity and performance, pushing the complexity limits further demands more sophisticated designs. In this letter, we propose CoFi-Lite, a highly efficient model that decouples spectral modeling into coarse- and fine-grained streams. By leveraging two parallel and symmetric encoder-decoder paths, it simultaneously extracts full-band envelopes and low-frequency details for complementary enhancement. In addition, a novel Cross-Path Fusion (CPF) module is introduced to bridge the distinct paths, facilitating efficient feature interaction. Remarkably, CoFi-Lite requires extremely low computational resources, featuring only 12.87M MACs/s and 83.12k parameters. Experimental results demonstrate that our proposed model outperforms the ultra-lightweight baseline GTCRN while requiring only 40.26% of its computational complexity. Its scaled-up variant also delivers performance on par with that of the SOTA ultra-lightweight model AdaptCRN alongside a 19.34% reduction in computational cost. Audio examples are available at https://acceleration123.github.io/CoFiLite-demo/.

eess.AS

Controlling radiative dynamics of a giant $\Lambda$-type atom via interference induced by the vacuum of a waveguide

We investigate the dynamics of a $\Lambda$-type giant atom (GA) whose both transition coupled to the guided modes of a one-dimensional (1D) waveguide at two spatially separated points with the GA initially excited and the electromagnetic (EM) modes of waveguide in vacuum. The spontaneous emission properties of this GA is investigated by solving the delay-differential equation for the amplitude of the 3GA in its excited state. Signatures of non-Markovian behavior is manifested in a population trapping in the excited state of the GA in the regime where the distance $d$ of the coupling points is smaller or comparable to the coherent length $L$ characterizing the width of the emitted wave packet. And an exact Markovian dynamics is also found when $d\geq L$ via the inference by adjusting the energy spacing and the inherent time delay besides the complex phases in the atom-light coupling, matching the behavior of a small atom coupled to a waveguide.

quant-ph

Pinning Down the Geometry of the Type Ic Broad-Line Supernova 2026gzf

Type Ic broad-line supernovae (SNe Ic-BL) are often associated with energetic explosions that display a prompt outburst of high-energy emission. Since their progenitor lost the H and He envelopes before the explosion exposing the C/O core, their explosion dynamics and geometry can be seen in an unobscured and undistorted way. We present imaging polarimetry and spectropolarimetry of the Type Ic-BL SN 2026gzf obtained 4.6 and 16.5 days after the X-ray shock breakout, which was recorded by the Einstein Probe satellite as EP260321a, showing it to be one of the softest and intrinsically dimmest extragalactic fast X-ray transients. The persistent low continuum polarization indicates that the outer layer of SN 2026gzf is mostly spherical, suggesting the explosion did not significantly disrupt the progenitor envelope. At day 16.5, the calcium near-infrared triplet displays a peak polarization above 1.5%. The geometry of the associated line opacity is also compatible with an axisymmetric configuration. The spatial distribution of such oxygen-burning ashes thus indicates the presence of a symmetry axis of the excitation structure within the nearly spherical ejecta. The Ca II triplet profile is dominated by a primary component spanning ~25,000--40,000 km/s, alongside a distinct secondary component extending above 28,000 km/s whose polarization implies a non-axisymmetric, complex excitation geometry toward the outer ejecta By implementing a three-dimensional Monte-Carlo calculation, we infer that a viewing angle of ~40 degree from the symmetry axis of the excitation structure could plausibly reproduce the observed spectral and polarization profiles of the Ca II triplet.

astro-ph.HE

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement

Flow matching (FM) enables high-fidelity generation, while self-supervised learning (SSL) speech models provide hierarchical representations spanning acoustic and phonetic levels. However, existing FM-based speech enhancement (SE) methods operate primarily in the spectral domain, treating SSL features only as external conditions rather than modeling directly in the SSL latent space. To fully exploit the structural richness of SSL representations, we propose PhASE-Flow, an FM-based SE framework that operates entirely in the SSL space. It models the conditional distribution of clean acoustic representations given phonetic ones, reconstructing the waveform via a neural vocoder. Experiments show that PhASE-Flow outperforms state-of-the-art baselines in perceptual quality and intelligibility. Notably, it achieves competitive performance with only four sampling steps, enabling highly efficient inference. Audio demos are available at https://anonymous.4open.science/w/phase-flow_demo-E6E1/.

eess.AS

HALO: Half-Frame-Rate Adaptive Learnable Operator for Lightweight STFT-Based Speech Enhancement

STFT-based speech enhancement typically adopts overlapping analysis frames. While overlap is essential for stable STFT processing, it makes adjacent frames highly correlated, causing redundant computation in lightweight models. We propose Half-frame-rate Adaptive Learnable Operator (HALO), a causal plug-in module that halves the internal frame rate without altering the STFT procedure. Broadly applicable to many lightweight models, HALO applies adaptive rate reduction before the backbone and restoration afterward, reconstructing the full-rate spectrum on the original STFT grid. Both reduction and restoration are implemented with lightweight dynamic convolutions. By halving the processed frame rate, HALO reduces backbone compute cost with no added algorithmic latency, freeing budget for channel widening. Experiments on the DNS3 dataset show consistent gains across diverse lightweight models under matched complexity, demonstrating the effectiveness of reducing overlap-induced redundancy.

eess.AS

FSC-Net: Integrating Fast Fourier Convolutions and Progressive Learning for Speech Bandwidth Extension

Speech bandwidth extension (BWE) aims to reconstruct high-fidelity wideband audio from narrowband inputs. While recent approaches have made significant progress, they often struggle to reconstruct realistic high-frequency phase and harmonic structures, leading to perceptual artifacts. In this paper, we propose FSC-Net (Full-Spectrum Context Network), a parameter-efficient architecture designed to explicitly model cross-band harmonic dependencies. By integrating Fast Fourier Convolutions (FFCs) into a complex spectral mapping framework, FSC-Net expands its receptive field to the entire spectrum, capturing long-range frequency interactions effectively. To address the ill-posed nature of high-frequency generation, our novel frequency-progressive learning curriculum guides the network to reconstruct spectral details from coarse to fine. Experimental results on the VCTK and unseen EARS datasets demonstrate that FSC-Net delivers consistently strong reconstruction quality and generalization, particularly in the challenging VCTK 4 kHz-to-48 kHz task. Compared to scaled-up baselines, our model attains leading LSD and PESQ scores while maintaining a highly compact parameter footprint (1.54 M).

eess.AS

Chips in the Flatland : 2D Semiconductors for Future Computing Electronic

As transistor scaling approaches its fundamental physical limits in the Angstrom era, two-dimensional (2D) semiconductors have emerged as the promising channel material candidates for future computing. While the device physics of 2D semiconductors have been rigorously explored, translating these nanodevices into fully functional integrated circuits remains a largely uncharted frontier. This review bridges the gap between material- and device-centric breakthroughs and circuit-level chip design in 2D semiconductors, a valley of death that has so far prevented translation of high-performance individual transistors into functional chips. We track the evolution of 2D semi-conductor field-effect transistors from basic Boolean logic families and standard cells to complex chip architectures, including recent milestones in RISC-V and monolithic CMOS microprocessors. Critically, we highlight the indispensable role of multiscale compact modeling, spanning semiclassical, quantum-hybrid and data-driven approaches, as the necessary link between device physics and the electronic design automation workflows for scalable chip development. By summarizing recent breakthroughs and identifying the bottlenecks in both fab and fabless trajectories of 2D semiconductors, this review shall provide insights that motivates the translation of proof-of-concept 2D transistors into fully functional computing chips, paving a way towards future Angstrom era computing technology empowered by 2D semiconductors.

physics.app-ph

Decoding Stimulus Reconstruction-Based Auditory Attention Robustly in Unbalanced EEG Datasets

In the past decade, numerous studies have applied deep neural networks (DNNs) to decode auditory attention (AAD) from Electroencephalogram (EEG) signals via stimulus reconstruction. However, the influence of dataset balance on the decoding performance of stimulus reconstruction-based AAD remains unexplored. In this study, three publicly available EEG-AAD datasets - KUL, DTU, and NJU cEEGrid - are used to construct both balanced and unbalanced experimental conditions. We hypothesize and demonstrate that stimulus reconstruction-based DNN decoders tend to produce overestimated decoding performance on unbalanced datasets. To address this issue, we propose a leave-one-paired-envelope-out (LOPEO) cross-validation protocol. Experimental results confirm that LOPEO effectively prevents inflated decoding accuracy on unbalanced datasets. While balanced datasets are generally preferred in experimental design, LOPEO provides a principled evaluation framework for unbalanced datasets that have already been published, filling an important gap in the field.

eess.AS

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptually plausible yet linguistically incorrect outputs. To address this issue, we propose L3-SE, a noise-invariant acoustic-semantic distillation framework for reducing linguistic hallucination in LM-based SE. The proposed method learns a noise-invariant conditioning encoder from noisy speech by jointly distilling two complementary clean-speech targets: an acoustic target for reconstruction fidelity and a semantic target for linguistic consistency. The resulting noise-invariant acoustic-semantic representations are used to condition a decoder-only autoregressive language model, which predicts clean acoustic tokens that are decoded into enhanced speech. To support high-quality generation, we further employ a high-fidelity codec built on learnable weighted WavLM layer representations as the discrete acoustic interface. By improving the reliability of conditioning under adverse conditions, the proposed framework substantially reduces hallucination and improves content faithfulness. Experiments show that the proposed method consistently outperforms prior LM-based speech enhancement baselines on linguistic consistency metrics, with especially clear gains under low-SNR and reverberant conditions, while maintaining competitive perceptual quality. Audio samples are available at https://max1wz.github.io/L3-SE-Demo-Page/. The complete source code will be released after the manuscript is accepted.

eess.AS

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations

Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is DeWavLM-Omni, a unified representation-level enhancement module fine-tuned from WavLM via knowledge distillation on a large-scale supervised multi-distortion dataset. This module directly converts degraded waveforms into clean and linguistically faithful phonetic representations, ensuring robust enhancement with minimal linguistic hallucination. Based on these enhanced phonetic representations, an Adapter generates enhanced acoustic representations containing rich acoustic details, which a neural Vocoder uses to reconstruct corresponding high-fidelity 16-kHz waveforms. A PostNet then converts the waveforms to 48~kHz before resampling them to their original rates, enabling seamless handling of inputs and outputs at multiple sampling rates. Experimental results on several evaluation datasets, covering sub-tasks and full tasks, demonstrate that UniPASE achieves superior or competitive performance compared with existing state-of-the-art models. The proposed model also serves as the backbone of our submission to the URGENT 2026 Challenge, which achieved 1st place in the objective evaluation. The source code and audio demos are available at https://github.com/xiaobin-rong/unipase/.

eess.AS

GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement

We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.

eess.AS

Traces of Helium Detected in Type Ic Supernova 2014L

The absence of helium features in optical spectra is one of the classification criteria for Type Ic supernovae (SNe Ic). However, it is highly debated whether helium is truly absent in ejecta or spectroscopically undetectable in the optical region. The near-infrared (NIR) region contains cleaner He lines that are less blended with other common ions in SNe Ic ejecta. We perform full spectral modeling on the near-peak-light optical and NIR spectra of the SN Ic 2014L to quantitatively constrain helium and other outer-ejecta properties, using the radiative transfer code TARDIS. We employ a deep-learning emulator for SNe Ic spectra that serves as a fast surrogate for TARDIS simulations. We then integrate the emulator within the Bayesian inference framework to infer the ejecta properties. The emulator achieves a mean fractional error of 1% between the emulated and TARDIS fluxes across all wavelengths and all samples in the test dataset. We constrain 0.018 to 0.020 M_sun (16% to 84% posterior percentile) of He above the photosphere near peak light in SN 2014L, inferred from the observed spectra covering 3500A to 24000A. A Bayesian statistical test shows that the observed spectra are inconsistent with no helium. Furthermore, the posterior favors a power-law density exponent of -7.04 to -6.88 (16% to 84% credible interval), consistent with theoretical calculations of radiation-dominated explosions. This work demonstrates that Bayesian radiative-transfer inference over a wide wavelength range provides a powerful path toward systematic constraints on He in SNe Ic.

astro-ph.HE

WhispSynth: Scaling Multilingual Whisper Corpus through Real Data Curation and A Novel Pitch-free Generative Framework

Whisper generation is constrained by the difficulty of data collection. Because whispered speech has low acoustic amplitude, high-fidelity recording is challenging. In this paper, we introduce WhispSynth, a large-scale multilingual corpus constructed via a novel high-fidelity generative framework. Specifically, we propose a pipeline integrating Differentiable Digital Signal Processing (DDSP)-based pitch-free method with Text-to-Speech (TTS) models. This framework refines a comprehensive collection of resources, including our newly constructed WhispNJU dataset, into 118 hours of high-fidelity whispered speech from 479 speakers. Unlike standard synthetic or noisy real data, our data engine faithfully preserves source vocal timbre and linguistic content while ensuring acoustic consistency, providing a robust foundation for text-to-whisper research. Experimental results demonstrate that WhispSynth exhibits significantly higher quality than existing corpora. Moreover, our CosyWhisper, tuned with WhispSynth, achieves speech naturalness on par with ground-truth samples. The official implementation and related resources are available at https://github.com/tan90xx/cosywhisper.

cs.SD

StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement

Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approach, PASE, is robust to hallucination but has limited perceptual quality under adverse conditions. We propose StuPASE, built upon PASE to achieve studio-level quality while retaining its low-hallucination property. First, we show that finetuning PASE with dry targets rather than targets containing simulated early reflections substantially improves dereverberation. Second, to address performance limitations under strong additive noise, we replace the GAN-based generative module in PASE with a flow-matching module, enabling studio-quality generation even under highly challenging conditions. Experiments demonstrate that StuPASE consistently produces perceptually high-quality speech while maintaining low hallucination, outperforming state-of-the-art SE methods. Audio demos are available at: https://xiaobin-rong.github.io/stupase_demo/.

eess.AS

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schr\"odinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.

eess.AS

Supernova Rates and Luminosity Functions from ASAS-SN III: Over a Decade of Type Ia SNe and Their Subtypes

We present volumetric rates and luminosity functions (LFs) of Type Ia supernovae (SNe Ia) from the All-Sky Automated Survey for Supernovae (ASAS-SN), covering the 11-year period from 2014 to 2024. By combining the 2014--2017 $V$-band sample with the 2018--2024 $g$-band sample, we construct a large statistical dataset of $1776$ SNe Ia. We compute completeness corrections based on injection-recovery simulations of the ASAS-SN light curves, taking into account the variations in light curve shapes. For our standard sample ($M_{g,\mathrm{peak}}<-16.0$ mag), we extract a total volumetric SN Ia rate of $R_{\mathrm{tot}} = (2.55 \pm 0.12) \times 10^4\,\mathrm{yr}^{-1}\,\mathrm{Gpc}^{-3}\,h_{70}^3$ at a median redshift of $z=0.029$. With a statistical uncertainty of $4.7\%$, this is the most precise local measurement to date. While the "normal" SNe Ia account for $(92.7 \pm 1.9)\%$ of this rate, the total LF reveals immense diversity, with $M_{g,\mathrm{peak}}$ spanning over five magnitudes. The LF of SNe Iax is also broad and rises toward lower luminosities, resulting in a likely lower limit of $(4.3 \pm 1.8)\%$ of the total rate. We place strong constraints on the rate of SNe Ia-CSM, finding they account for only $(0.036 \pm 0.017)\%$ of the total local rate. Finally, we find that the low-luminosity 02es-like SNe are $7 \pm 5$ times more common than the luminous 03fg-like SNe. This places demographic constraints on models proposing a physical continuum for these two subtypes, implying that any common channel for the two classes must strongly favor lower-luminosity explosions.

astro-ph.HE