SearcharxivSearch

arXiv subjects

Zhu Liu

Publications and source records attributed to Zhu Liu.

At least 19 recordsLinked to original sources

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme level remains challenging. We propose CtrlSpeech, a controllable, expressive TTS framework with coarse-to-fine control. Built on the DiTAR architecture, CtrlSpeech combines global speaker conditioning with phone-aligned pitch, loudness, and duration signals, enabling localized prosodic control while preserving the target speaker's timbre. This design allows users to adjust expressive attributes at a fine temporal granularity, making speech refinement more flexible and controllable. Experimental results show that CtrlSpeech achieves competitive zero-shot TTS performance and improves controllability over expressive attributes, demonstrating its effectiveness for flexible and practical expressive speech synthesis.

eess.AS

AugCodec: A Low-Bitrate Disentangled Neural Speech Codec via Data Augmentation

We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and prosody tokens. Specifically, we employ tailored augmenta tion strategies to transform speech into distinct variants, each serving as input for extracting tokens that preserve the target attribute while suppressing others. This disentanglement strategy enables substantial reduction in token rate. Further more, we introduce an augmentation loss that aligns semantic encoder outputs between source and voice-converted speech, encouraging speaker-agnostic embeddings while mitigating the acoustic mismatch induced by voice conversion. Experiments on LibriSpeech test-clean demonstrate that AugCodec significantly outperforms state-of-the-art methods in both reconstruction quality and disentanglement, while operating at only 12.5Hz with three token streams.

cs.SD

Unveiling Orbital-mediated Ultrafast Demagnetization in Rare Earth-Transition-Metal Ferrimagnets

The ultimate speed limit of magnetic recording and spintronic devices is set by the efficiency of angular-momentum transfer during ultrafast demagnetization, yet its microscopic pathway in Rare-Earth-Transition-Metal (RE-TM) ferrimagnets remains debated. Here, we establish an orbital-mediated framework in which 3d spin-orbit coupling (SOC) governs angular momentum (AM) dissipation. Strong 3d-SOC in RE-Co enables sub-picosecond, single-step demagnetization via direct orbital-to-lattice transfer, whereas weak 3d-SOC in RE-Fe redirects AM into 4f orbitals, producing slower two-step dynamics. The second-stage rate scales with 4f-SOC strength, revealing a distinct orbital-mediated dissipation channel. Using time-resolved magneto-optical Kerr measurements, supported by an extended four-temperature model, corroborate this picture across diverse RE-TM systems (RE = Sm, Gd, Tb, Dy, Ho and TM = Fe, Co, CoNi). Our results identify the SOC-driven competition between 3d and 4f orbital channels as the universal mechanism governing ultrafast demagnetization in RE-TM ferrimagnets, enabling rational design of the switching speed for next-generation spintronic devices.

cond-mat.mtrl-sci

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo-label evolution in cluttered, low-contrast infrared imagery and severe sample-distribution imbalance. In this paper, we present a more adaptive and stable framework to address these issues. Leveraging the intrinsic consistency between thermal radiation patterns and heat diffusion, we propose a physics-induced annotation strategy that expands single-point labels into reliable pseudo-masks. To further enhance supervision and alleviate sample imbalance, we develop a bi-level dual-update framework that jointly optimizes detector weights, sample weights, and diffusion parameters. A meta-classifier dynamically predicts sample-wise loss weights, while a differentiable diffusion module refines pseudo-labels with detection feedback, enabling adaptive interaction between training and hyperparameter optimization. Extensive experiments across multiple datasets demonstrate five-fold annotation acceleration, superior detection accuracy, and comparable performance with 30% of the training data, validating the efficiency and practicality of our approach. Our code is available at https://github.com/yuanhang-yao/diffuse-to-detect.

cs.CV

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics, leading to noisy pseudo-masks and unstable optimization. To address this, we propose a hierarchical VFM-driven knowledge distillation framework that uses a frozen Vision Foundation Model (VFM) during training. We formulate point-supervised learning as a bilevel optimization process: the inner loop adapts a VFM-embedded teacher on reweighted training samples, while the outer loop transfers validation-guided knowledge to a lightweight student to mitigate pseudo-label noise and training-set bias. We further introduce Semantic-Conditioned Affine Modulation (SCAM) to inject VFM semantics into CNN features at multiple layers. In addition, a dynamic collaborative learning strategy with cluster-level sample reweighting enhances robustness to imperfect pseudo-masks. Experiments on diverse challenging cases across multiple ISTD backbones demonstrate consistent improvements in detection accuracy and training stability. Our code is available at https://github.com/yuanhang-yao/semantic-prior.

cs.CV

Labeled TrustSet Guided: Batch Active Learning with Reinforcement Learning

Batch active learning (BAL) is a crucial technique for reducing labeling costs and improving data efficiency in training large-scale deep learning models. Traditional BAL methods often rely on metrics like Mahalanobis Distance to balance uncertainty and diversity when selecting data for annotation. However, these methods predominantly focus on the distribution of unlabeled data and fail to leverage feedback from labeled data or the model's performance. To address these limitations, we introduce TrustSet, a novel approach that selects the most informative data from the labeled dataset, ensuring a balanced class distribution to mitigate the long-tail problem. Unlike CoreSet, which focuses on maintaining the overall data distribution, TrustSet optimizes the model's performance by pruning redundant data and using label information to refine the selection process. To extend the benefits of TrustSet to the unlabeled pool, we propose a reinforcement learning (RL)-based sampling policy that approximates the selection of high-quality TrustSet candidates from the unlabeled data. Combining TrustSet and RL, we introduce the Batch Reinforcement Active Learning with TrustSet (BRAL-T) framework. BRAL-T achieves state-of-the-art results across 10 image classification benchmarks and 2 active fine-tuning tasks, demonstrating its effectiveness and efficiency in various domains.

cs.LG

Global near-real-time daily emissions of atmospheric pollutants from power plants

The power sector is a major source of fossil fuel use and air pollutant emissions, making high-spatiotemporal-resolution emission accounting essential for effective mitigation policy and air quality management. Yet existing public inventories are often limited by low timeliness and coarse resolution. Here, we develop a global, plant-level, daily, multi-pollutant emission database for the power sector by integrating nearly 3 million hourly-to-daily near-real-time power generation records from 57 countries, representing about 81% of global fossil-fuel-based electricity generation, with fundamental information for more than 10,000 power plants worldwide, including location and installed capacity. The dataset substantially improves the timeliness and granularity of global power-sector emission estimates. From 2019 to 2025, emissions of most pollutants increased, with 2025 daily mean emissions reaching 0.274 kt/d for BC, 45.1 kt/d for CO, 0.418 kt/d for NH3, 52.2 kt/d for NOx, 3.01 kt/d for NMVOC, 0.418 kt/d for OC, 6.76 kt/d for PM10, 5.11 kt/d for PM2.5, and 78.5 kt/d for SO2. Compared with 2019, NMVOC showed the largest increase, whereas SO2 was the only pollutant to decline overall. Coal remained the dominant source of sulfur-, nitrogen-, and particulate-related emissions, while gas and biomass contributed more to carbonaceous species and reduced nitrogen. The dataset also captures pronounced seasonal, regional, and short-term variability. Against EDGAR for 2019-2022, our estimates agree well, with Pearson correlations of 0.92-0.99 and mean relative deviations of 8.8%-28.1%. This near-real-time, high-resolution dataset provides a strong foundation for air pollution control, carbon mitigation, emission monitoring, and satellite-based inversion.

physics.ao-ph

Near real-time monitoring of global land-ocean cover dynamics

Monitoring the dynamics of global land-ocean cover is fundamental for regulating the Earth's climate and sustaining terrestrial and marine ecosystems. However, existing datasets and research often exhibit limitations in temporal resolution and timeliness, lack coupled analysis of land cover and sea ice dynamics, and fail to incorporate the perspective of Earth system safety thresholds. Here, we developed an integrated monitoring framework by fusing multi-source remote sensing and reanalysis data, generating a 5-day resolution time series (2018-2025) of global land cover and sea ice coverage with near-real-time update capability. Our analysis reveals distinct latitudinal and regional patterns, with forests dominating (27.0% of global land area) tropical and subtropical regions. At the national scale, land cover composition and seasonal rhythms vary significantly, with countries like China, India, and the US exhibiting divergent patterns such as bimodal cropland fluctuations and alternating snow/ice dominance. Temporally, vegetated cover types exhibit seasonal cycles peaking during Northern Hemisphere summer, and a pronounced anti-phase seasonal pattern is observed between Arctic and Antarctic sea ice coverage. Crucially, safety threshold analysis indicates the global forest cover indicator (~60%) is approaching the 54% lower safe limit, with a declining trend in recent years. Concurrently, Arctic sea ice coverage in September occasionally drops to 23%, below its critical upper limit of 27.6%. Temperature presents a significant negative correlation with sea ice cover (R = -0.78, p < 0.001), with asymmetric freezing and melting rates. By quantifying the proximity of key indicators to their safety thresholds, this study provides a robust, integrated framework for early-warning assessment, thereby offering vital scientific support for global climate adaptation and sustainable policymaking.

physics.ao-ph

CTFS : Collaborative Teacher Framework for Forward-Looking Sonar Image Semantic Segmentation with Extremely Limited Labels

As one of the most important underwater sensing technologies, forward-looking sonar exhibits unique imaging characteristics. Sonar images are often affected by severe speckle noise, low texture contrast, acoustic shadows, and geometric distortions. These factors make it difficult for traditional teacher-student frameworks to achieve satisfactory performance in sonar semantic segmentation tasks under extremely limited labeled data conditions. To address this issue, we propose a Collaborative Teacher Semantic Segmentation Framework for forward-looking sonar images. This framework introduces a multi-teacher collaborative mechanism composed of one general teacher and multiple sonar-specific teachers. By adopting a multi-teacher alternating guidance strategy, the student model can learn general semantic representations while simultaneously capturing the unique characteristics of sonar images, thereby achieving more comprehensive and robust feature modeling. Considering the challenges of sonar images, which can lead teachers to generate a large number of noisy pseudo-labels, we further design a cross-teacher reliability assessment mechanism. This mechanism dynamically quantifies the reliability of pseudo-labels by evaluating the consistency and stability of predictions across multiple views and multiple teachers, thereby mitigating the negative impact caused by noisy pseudo-labels. Notably, on the FLSMD dataset, when only 2% of the data is labeled, our method achieves a 5.08% improvement in mIoU compared to other state-of-the-art approaches.

cs.CV

Discovery of crested quasi-periodic eruptions following the most luminous SRG/eROSITA tidal disruption event

We report the discovery of complex flaring activity from the galactic nucleus hosting the five-year-old tidal disruption event eRASSt J234402.9-352640 (J2344). With Einstein Probe and XMM-Newton observations, we detected highly structured soft X-ray variability. Through temporal decomposition of the XMM-Newton light curve and time-resolved spectral analysis, we identified broad, thermal flares recurring every $\sim$12 hours and lasting $\sim$2 hours, consistent with quasi-periodic eruptions (QPEs). Remarkably, these QPEs are accompanied by an unprecedented crest of hotter, shorter flares, each lasting between 5 and 30 minutes. These flares are predominantly found in the rising phases of the QPEs, although they also appear throughout the quiescence. These findings establish J2344 as a new member of the QPE emitter population and uncover a previously unobserved phenomenology that challenges current models of QPEs. In this letter, we present the phenomenological properties of this unique source and discuss possible interpretations within the framework of extreme-mass-ratio inspirals.

astro-ph.HE

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their inability to provide precise annotation data for sonar images. Therefore, designing effective object detection methods for sonar images with extremely limited labels is particularly important. To address this, we propose a teacher-student framework called RSOD, which aims to fully learn the characteristics of sonar images and develop a pseudo-label strategy suitable for these images to mitigate the impact of limited labels. First, RSOD calculates a reliability score by assessing the consistency of the teacher's predictions across different views. To leverage this score, we introduce an object mixed pseudo-label method to tackle the shortage of labeled data in sonar images. Finally, we optimize the performance of the student by implementing a reliability-guided adaptive constraint. By taking full advantage of unlabeled data, the student can perform well even in situations with extremely limited labels. Notably, on the UATD dataset, our method, using only 5% of labeled data, achieves results that can compete against those of our baseline algorithm trained on 100% labeled data. We also collected a new dataset to provide more valuable data for research in the field of sonar.

cs.CV

RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech

End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we introduce RosettaSpeech, a novel zero-shot framework trained exclusively on monolingual speech-text data augmented by machine translation supervision. Unlike prior works that rely on complex cascaded pseudo-labeling, our approach strategically utilizes text as a semantic bridge during training to synthesize translation targets, thereby eliminating the need for parallel speech pairs while maintaining a direct, end-to-end inference pipeline. Empirical evaluations on the CVSS-C benchmark demonstrate that RosettaSpeech achieves state-of-the-art zero-shot performance, surpassing leading baselines by significant margins - achieving ASR-BLEU scores of 25.17 for German-to-English (+27% relative gain) and 29.86 for Spanish-to-English (+14%). Crucially, our model effectively preserves the source speaker's voice without ever seeing paired speech data. We further analyze the impact of data scaling and demonstrate the model's capability in many-to-one translation, offering a scalable solution for extending high-quality S2ST to "text-rich, speech-poor" languages.

eess.AS

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, German, Dutch, Italian, Portuguese, and Polish. VoiceCraft-X utilizes the Qwen3 large language model for phoneme-free cross-lingual text processing and a novel token reordering mechanism with time-aligned text and speech tokens to handle both tasks as a single sequence generation problem. The model generates high-quality, natural-sounding speech, seamlessly creating new audio or editing existing recordings within one framework. VoiceCraft-X shows robust performance in diverse linguistic settings, even with limited per-language data, underscoring the power of unified autoregressive approaches for advancing complex, real-world multilingual speech applications. Audio samples are available at https://zhishengzheng.com/voicecraft-x/.

eess.AS

EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query localization in egocentric vision. Inspired by avian memory consolidation, EAGLE synergistically integrates segmentation guided by an appearance-aware meta-learning memory (AMM), with tracking driven by a geometry-aware localization memory (GLM). This memory consolidation mechanism, through structured appearance and geometry memory banks, stores high-confidence retrieval samples, effectively supporting both long- and short-term modeling of target appearance variations. This enables precise contour delineation with robust spatial discrimination, leading to significantly improved retrieval accuracy. Furthermore, by integrating the VQL-2D output with a visual geometry grounded Transformer (VGGT), we achieve a efficient unification of 2D and 3D tasks, enabling rapid and accurate back-projection into 3D space. Our method achieves state-ofthe-art performance on the Ego4D-VQ benchmark.

cs.CV

ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement

Coreference Resolution (CR) is a critical task in Natural Language Processing (NLP). Current research faces a key dilemma: whether to further explore the potential of supervised neural methods based on small language models, whose detect-then-cluster pipeline still delivers top performance, or embrace the powerful capabilities of Large Language Models (LLMs). However, effectively combining their strengths remains underexplored. To this end, we propose \textbf{ImCoref-CeS}, a novel framework that integrates an enhanced supervised model with LLM-based reasoning. First, we present an improved CR method (\textbf{ImCoref}) to push the performance boundaries of the supervised neural method by introducing a lightweight bridging module to enhance long-text encoding capability, devising a biaffine scorer to comprehensively capture positional information, and invoking a hybrid mention regularization to improve training efficiency. Importantly, we employ an LLM acting as a multi-role Checker-Splitter agent to validate candidate mentions (filtering out invalid ones) and coreference results (splitting erroneous clusters) predicted by ImCoref. Extensive experiments demonstrate the effectiveness of ImCoref-CeS, which achieves superior performance compared to existing state-of-the-art (SOTA) methods.

cs.CL

Enhancing Infrared Vision: Progressive Prompt Fusion Network and Benchmark

We engage in the relatively underexplored task named thermal infrared image enhancement. Existing infrared image enhancement methods primarily focus on tackling individual degradations, such as noise, contrast, and blurring, making it difficult to handle coupled degradations. Meanwhile, all-in-one enhancement methods, commonly applied to RGB sensors, often demonstrate limited effectiveness due to the significant differences in imaging models. In sight of this, we first revisit the imaging mechanism and introduce a Progressive Prompt Fusion Network (PPFN). Specifically, the PPFN initially establishes prompt pairs based on the thermal imaging process. For each type of degradation, we fuse the corresponding prompt pairs to modulate the model's features, providing adaptive guidance that enables the model to better address specific degradations under single or multiple conditions. In addition, a Selective Progressive Training (SPT) mechanism is introduced to gradually refine the model's handling of composite cases to align the enhancement process, which not only allows the model to remove camera noise and retain key structural details, but also enhancing the overall contrast of the thermal image. Furthermore, we introduce the most high-quality, multi-scenarios infrared benchmark covering a wide range of scenarios. Extensive experiments substantiate that our approach not only delivers promising visual results under specific degradation but also significantly improves performance on complex degradation scenes, achieving a notable 8.76\% improvement. Code is available at https://github.com/Zihang-Chen/HM-TIR.

cs.CV

4D-PreNet: A Unified Preprocessing Framework for 4D-STEM Data Analysis

Automated experimentation with real time data analysis in scanning transmission electron microscopy (STEM) often require end-to-end framework. The four-dimensional scanning transmission electron microscopy (4D-STEM) with high-throughput data acquisition has been constrained by the critical bottleneck results from data preprocessing. Pervasive noise, beam center drift, and elliptical distortions during high-throughput acquisition inevitably corrupt diffraction patterns, systematically biasing quantitative measurements. Yet, conventional correction algorithms are often material-specific and fail to provide a robust, generalizable solution. In this work, we present 4D-PreNet, an end-to-end deep-learning pipeline that integrates attention-enhanced U-Net and ResNet architectures to simultaneously perform denoising, center correction, and elliptical distortion calibration. The network is trained on large, simulated datasets encompassing a wide range of noise levels, drift magnitudes, and distortion types, enabling it to generalize effectively to experimental data acquired under varying conditions. Quantitative evaluations demonstrate that our pipeline reduces mean squared error by up to 50% during denoising and achieves sub-pixel center localization in the center detection task, with average errors below 0.04 pixels. The outputs are bench-marked against traditional algorithms, highlighting improvements in both noise suppression and restoration of diffraction patterns, thereby facilitating high-throughput, reliable 4D-STEM real-time analysis for automated characterization.

cs.CV

A new Bowen Fluorescence Flare and Extreme Coronal Line Emitter discovered by SRG/eROSITA

The nuclear transient eRASSt J012026.5-292727 (J012026 hereafter) was discovered in the second SRG/eROSITA all-sky survey (eRASS2). The source appeared more than one order of magnitude brighter than the eRASS1 upper limits (peak eRASS2 0.2-2.3 keV flux of 1.14 x 10^-12 erg cm^-2 s^-1), and with a soft X-ray spectrum (photon index Gamma = 4.3). Over the following months, the X-ray flux started decaying, with significant flaring activity on both hour- and year-timescales. By inspecting the multiwavelength light curves of time-domain wide-field facilities, we detected a strong mid-infrared flare, evolving over 2 years, and a weaker optical counterpart. Follow-up optical spectroscopy revealed transient features, including redshifted Balmer lines (FWHM ~1500 km/s), strong Fe II emission, He II and Bowen lines, and high-ionization iron coronal lines. One spectrum showed a triple-peaked H-beta line, consistent with emission from a face-on elliptical disk. The spectroscopic features and the slow evolution of the event place J012026 within the classifications of Bowen fluorescence flares (BFFs) and extreme coronal line emitters (ECLEs). BFFs have been associated with rejuvenated accreting SMBHs, although the mechanism triggering the onset of the new accretion flow is still unclear, while ECLEs have been linked to the disruption of stars in gas-rich environments. The association of J012026 to both classes, combined with the multi-wavelength information, suggests that BFFs could be, at least in some cases, due to tidal disruption events (TDEs). The observed X-ray variability, uncommon in standard TDEs, adds complexity to these families of nuclear transients. These results highlight the diverse phenomenology of nuclear accretion events and demonstrate the value of systematic X-ray surveys, such as eROSITA and Einstein Probe, for uncovering such transients and characterizing their physical origin.

astro-ph.HE