SearcharxivSearch

arXiv subjects

Yi Chai

Publications and source records attributed to Yi Chai.

14 recordsLinked to original sources

RegimeFormer: A Large Protein Model of Global Perturbation Regimes

Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.

q-bio.QM

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable contrasts and restoring degraded acquisitions. Across ten completion tasks, task-specific MSCNet achieved mean structural similarity of 0.818 versus 0.798 for the strongest task-matched comparators; matched-capacity analyses showed larger differences in lesion fidelity and boundary preservation. In a blinded 1,000-case reader study, overall image quality met the prespecified non-inferiority criterion for DWI, ADC and T2W completion, but not T1W. In a separate 200-case diagnostic assessment, AUCs for clinically significant cancer were 0.860 with acquired images, 0.841 with MSCNet and 0.797 with baseline-generated images. A locked 186-case three-hospital cohort supported multicentre transportability. These retrospective results support quality-controlled cross-modal reconstruction as an adjunct to acquired prostate MRI.

eess.IV

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.

cs.SD

Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains unclear whether indiscriminate scaling commensurately improves model generalization. This study challenges the "scale-first" paradigm by decoupling the impacts of training data scale versus diversity. Through experiments on representative datasets, we report two key findings: (1) Larger is not always better. Expanding data scale excessively under fixed generation methods yields negligible returns and may even degrade cross-domain generalization due to overfitting.(2) Diversity outweighs scale. A smaller composite training set featuring diverse attacks significantly outperforms larger-scale datasets with limited diversity in cross-dataset evaluations. We conclude that future dataset construction should prioritize the diversity of generation methods over scale to effectively enhance model generalization.

cs.SD

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI) such as public figures. Current detection systems primarily rely on generic, black-box models that fail to capture speaker-specific idiosyncratic traits and lack interpretability. In this paper, we propose Phoneme-based Voice Profiling (PVP), a novel personalized defense framework. By shifting the detection paradigm from macro-utterance analysis to micro-phonetic modeling, PVP captures the unique acoustic distributions underlying a POI's habitual articulatory patterns. Specifically, our framework models speaker-specific phonetic realizations using lightweight Gaussian Mixture Models (GMMs) estimated solely from bona fide reference speech. This design enables data-efficient profiling and robust generalization to previously unseen spoofing attacks without requiring heavy spoof-specific training. Furthermore, we introduce the first large-scale Chinese POI deepfake dataset to benchmark speaker-specific detection. Experimental results demonstrate that PVP significantly outperforms state-of-the-art generic detectors in POI spoofing scenarios, achieving substantial EER reductions while providing fine-grained, phoneme-level interpretability for forensic analysis. Code and data are available at: https://github.com/JunXue-tech/PVP

cs.SD

How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?

Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and transmission mechanisms on social media platforms further obscure deepfake artifacts. These factors complicate reliable detection in real-world environments, underscoring the need for representative evaluation benchmarks. To this end, we introduce ML-ITW (Multilingual In-The-Wild), a multilingual dataset covering 14 languages, seven major platforms, and 180 public figures, totaling 28.39 hours of audio. We evaluate three detection paradigms: end-to-end neural models, self-supervised feature-based (SSL) methods, and audio large language models (Audio LLMs). Experimental results reveal significant performance degradation across diverse languages and real-world acoustic conditions, highlighting the limited generalization ability of existing detectors in practical scenarios. The ML-ITW dataset is publicly available.

cs.SD

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting in restricted diversity and poor coverage of realistic editing scenarios. Meanwhile, current SED methods rely heavily on frame-level supervision to detect observable acoustic anomalies, which fundamentally limits their ability to handle deletion-type edits, where the manipulated content is entirely absent from the signal. To address these challenges, we present a unified framework that bridges speech editing detection and content localization through a generative formulation based on Audio Large Language Models (Audio LLMs). We first introduce AiEdit, https://huggingface.co/datasets/JunXueTech/AiEdit, a large-scale bilingual dataset (approximately 140 hours) that covers addition, deletion, and modification operations using state-of-the-art end-to-end speech editing systems, providing a more realistic benchmark for modern threats. Building upon this, we reformulate SED as a structured text generation task, enabling joint reasoning over edit type identification, and content localization. To enhance the grounding of generative models in acoustic evidence, we propose a prior-enhanced prompting strategy that injects word-level probabilistic cues derived from a frame-level detector. Furthermore, we introduce an acoustic consistency-aware loss that explicitly enforces the separation between normal and anomalous acoustic representations in the latent space. Experimental results demonstrate that the proposed approach consistently outperforms existing methods across both detection and localization tasks.

cs.SD

Audio-visual Event Localization on Portrait Mode Short Videos

Audio-visual event localization (AVEL) plays a critical role in multimodal scene understanding. While existing datasets for AVEL predominantly comprise landscape-oriented long videos with clean and simple audio context, short videos have become the primary format of online video content due to the the proliferation of smartphones. Short videos are characterized by portrait-oriented framing and layered audio compositions (e.g., overlapping sound effects, voiceovers, and music), which brings unique challenges unaddressed by conventional methods. To this end, we introduce AVE-PM, the first AVEL dataset specifically designed for portrait mode short videos, comprising 25,335 clips that span 86 fine-grained categories with frame-level annotations. Beyond dataset creation, our empirical analysis shows that state-of-the-art AVEL methods suffer an average 18.66% performance drop during cross-mode evaluation. Further analysis reveals two key challenges of different video formats: 1) spatial bias from portrait-oriented framing introduces distinct domain priors, and 2) noisy audio composition compromise the reliability of audio modality. To address these issues, we investigate optimal preprocessing recipes and the impact of background music for AVEL on portrait mode videos. Experiments show that these methods can still benefit from tailored preprocessing and specialized model design, thus achieving improved performance. This work provides both a foundational benchmark and actionable insights for advancing AVEL research in the era of mobile-centric video content. Dataset and code will be released.

cs.MM

Atacama Large Aperture Submillimeter Telescope (AtLAST) Science: Solar and stellar observations

Observations at (sub-)millimeter wavelengths offer a complementary perspective on our Sun and other stars, offering significant insights into both the thermal and magnetic composition of their chromospheres. Despite the fundamental progress in (sub-)millimeter observations of the Sun, some important aspects require diagnostic capabilities that are not offered by existing observatories. In particular, simultaneous observations of the radiation continuum across an extended frequency range would facilitate the mapping of different layers and thus ultimately the 3D structure of the solar atmosphere. Mapping large regions on the Sun or even the whole solar disk at a very high temporal cadence would be crucial for systematically detecting and following the temporal evolution of flares, while synoptic observations, i.e., daily maps, over periods of years would provide an unprecedented view of the solar activity cycle in this wavelength regime. As our Sun is a fundamental reference for studying the atmospheres of active main sequence stars, observing the Sun and other stars with the same instrument would unlock the enormous diagnostic potential for understanding stellar activity and its impact on exoplanets. The Atacama Large Aperture Submillimeter Telescope (AtLAST), a single-dish telescope with 50\,m aperture proposed to be built in the Atacama desert in Chile, would be able to provide these observational capabilities. Equipped with a large number of detector elements for probing the radiation continuum across a wide frequency range, AtLAST would address a wide range of scientific topics including the thermal structure and heating of the solar chromosphere, flares and prominences, and the solar activity cycle. In this white paper, the key science cases and their technical requirements for AtLAST are discussed.

astro-ph.SR

Evaluating Non-LTE Spectral Inversions with ALMA and IBIS

We present observations of a solar plage in the millimeter-continuum with the ALMA and in the Ca 8542 and Na 5896 spectral lines with the Interferometric BIdimensional Spectrometer (IBIS). Our goal is to compare the measurement of local gas temperatures provided by ALMA with the temperature diagnostics provided by non-LTE inversions using STIC. In performing these inversions, we find that using column mass as the reference height scale, rather than optical depth, provides more reliable atmospheric profiles above the temperature minimum and that the treatment of non- LTE hydrogen ionization brings the inferred chromospheric temperatures into better agreement with the ALMA measurements. The Band 3 brightness temperatures are higher but well correlated with the inversion-derived temperatures at the height of formation of the Ca 8542 line core. The Band 6 temperatures instead do not show good correlations with the temperatures at any specific layer in the inverted atmospheres. We then performed inversions that included the millimeter continuum intensities as an additional constraint. Incorporating Band 3 generally resulted in atmospheres showing a strong temperature rise in the upper atmosphere, while including Band 6 led to significant regions of anomalously low temperatures at chromospheric heights. This is consistent with the idea that the Band 6 emission can come from a range of heights. The poor constraints on the chromospheric electron density with existing inversion codes introduces difficulties in determining the height(s) of formation of the millimeter continuum as well as uncertainties in the temperatures derived from the spectral lines.

astro-ph.SR

A study of sunspot 3 minute oscillations using ALMA and GST

Waves and oscillations are important solar phenomena, not only because they can propagate and dissipate energy in the chromosphere, but also because they carry information about the structure of the atmosphere in which they propagate. The nature of the three-minute oscillations observed in the umbral region of sunspots is considered to be an effect of propagation of magnetohydrodynamic (MHD) waves upward from below the photosphere. We present a study of sunspot oscillations and wave propagation in NOAA AR 12470 using an approximately one-hour long data set acquired on 2015 December 17 by the Atacama Large Millimeter/submillimeter Array (ALMA), the Goode Solar Telescope (GST) operating at the Big Bear Solar Observatory (BBSO), the Atmospheric Imaging Assembly (AIA) on board the Solar Dynamics Observatory (SDO), and the Interface Region Imaging Spectrograph (IRIS). The ALMA data are unique in providing a time-series of direct temperature measurements in the sunspot chromosphere. The two-second cadence of ALMA images allows us to well resolve the three-minute periods typical of sunspot oscillations in the chromosphere. Fourier analysis is applied to ALMA Band 3 ($\sim$100 GHz, $\sim$3 mm) and GST H$α$ data sets to obtain power spectra as well as oscillation phase information. We analysed properties of the wave propagation by combining multiple wavelengths that probe physical parameters of solar atmosphere at different heights. We find that the ALMA temperature fluctuations are consistent with that expected for a propagating acoustic wave, with a slight asymmetry indicating non-linear steepening.

astro-ph.SR

High-frequency wave power observed in the solar chromosphere with IBIS and ALMA

We present observational constraints on the solar chromospheric heating contribution from acoustic waves with frequencies between 5 and 50 mHz. We utilize observations from the Dunn Solar Telescope in New Mexico complemented with observations from the Atacama Large Millimeter Array collected on 2017 April 23. The properties of the power spectra of the various quantities are derived from the spectral lines of Ca II 854.2 nm, H I 656.3 nm, and the millimeter continuum at 1.25 mm and 3 mm. At the observed frequencies the diagnostics almost all show a power law behavior, whose particulars (slope, peak and white noise floors) are correlated with the type of solar feature (internetwork, network, plage). In order to disentangle the vertical versus transverse plasma motions we examine two different fields of view; one near disk center and the other close to the limb. To infer the acoustic flux in the middle chromosphere, we compare our observations with synthetic observables from the time-dependent radiative hydrodynamic RADYN code. Our findings show that acoustic waves carry up to about 1 kW m$^{-2}$ of energy flux in the middle chromosphere, which is not enough to maintain the quiet chromosphere, contrary to previous publications.

astro-ph.SR

Solar Chromospheric Temperature Diagnostics: a joint ALMA-H$α$ analysis

We present the first high-resolution, simultaneous observations of the solar chromosphere in the optical and millimeter wavelength ranges, obtained with ALMA and the IBIS instrument at the Dunn Solar Telescope. In this paper we concentrate on the comparison between the brightness temperature observed in ALMA Band 3 (3 mm; 100 GHz) and the core width of the H$α$ 656.3 nm line, previously identified as a possible diagnostic of the chromospheric temperature. We find that in the area of plage, network and fibrils covered by our FOV the two diagnostics are well correlated, with similar spatial structures observed in both. The strength of the correlation is remarkable, given that the source function of the mm-radiation obeys local thermodynamic equilibrium, while the H$α$ line has a source function that deviates significantly from the local Planck function. The observed range of ALMA brightness temperatures is sensibly smaller than the temperature range that was previously invoked to explain the observed width variations in H$α$. We employ analysis from forward modeling with the RH code to argue that the strong correlation between H$α$ width and ALMA brightness temperature is caused by their shared dependence on the population number $n_2$ of the first excited level of hydrogen. This population number drives millimeter opacity through hydrogen ionization via the Balmer continuum, and H$α$ width through a curve-of-growth-like opacity effect. Ultimately, the $n_2$ population is regulated by the enhancement or lack of downward Ly$α$ flux, which coherently shifts the formation height of both diagnostics to regions with different temperature, respectively.

astro-ph.SR

Forecasting Spatio-Temporal Renewable Scenarios: a Deep Generative Approach

The operation and planning of large-scale power systems are becoming more challenging with the increasing penetration of stochastic renewable generation. In order to minimize the decision risks in power systems with large amount of renewable resources, there is a growing need to model the short-term generation uncertainty. By producing a group of possible future realizations for certain set of renewable generation plants, scenario approach has become one popular way for renewables uncertainty modeling. However, due to the complex spatial and temporal correlations underlying in renewable generations, traditional model-based approaches for forecasting future scenarios often require extensive knowledge, while fitted models are often hard to scale. To address such modeling burdens, we propose a learning-based, data-driven scenario forecasts method based on generative adversarial networks (GANs), which is a class of deep-learning generative algorithms used for modeling unknown distributions. We firstly utilize an improved GANs with convergence guarantees to learn the intrinsic patterns and model the unknown distributions of (multiple-site) renewable generation time-series. Then by solving an optimization problem, we are able to generate forecasted scenarios without any scenario number and forecasting horizon restrictions. Our method is totally model-free, and could forecast scenarios under different level of forecast uncertainties. Extensive numerical simulations using real-world data from NREL wind and solar integration datasets validate the performance of proposed method in forecasting both wind and solar power scenarios.

math.OC