SearcharxivSearch

arXiv subjects

Zhao Guo

Publications and source records attributed to Zhao Guo.

At least 19 recordsLinked to original sources

Tidal dissipation in magnetised, rotating stars and planets: linear calculations exploring various magnetic field configurations

We study tidal flows in the convective envelopes of rotating, magnetised fluid bodies, such as low-mass stars and giant planets. In well-mixed convective regions, (magneto-)inertial waves are linearly excited by tidal forcing, and their dissipation can dominantly drive spin and orbital evolution in many close star-planet and binary star systems. We perform linear magnetohydrodynamic calculations of wavelike tides in spherical-shell geometry of a tidally-forced, rotating, incompressible, viscous and non-ideal magnetised fluid. Our calculations consider the widest range of magnetic field configurations to date (including both aligned and misaligned dipole fields, free-decay dipole and quadrupole fields, azimuthal "Malkus fields" and mixed poloidal-toroidal "Prendergast fields") to analyse the effects of magnetic fields on the wavelike response and dissipation. We find that the tidal response at a given frequency depends strongly on both magnetic field strength and geometry. Magnetic fields with strong poloidal components modify the flow more efficiently and introduce high-frequency Alfv\'enic resonances associated with weakly damped eigenmodes. When an enhanced (turbulent) viscosity is adopted, we find that viscous dissipation remains comparable to Ohmic dissipation for strong fields, in contrast to previous studies in which Ohmic dissipation was argued to dominate. We also explore the variation in magnetic effects as the shell thickness, magnetic Prandtl and Ekman numbers are varied. Finally, the frequency-averaged tidal power is found to be largely insensitive to the magnetic field in most cases, though significant deviations are found for free-decay fields. Our results have important implications for the tidal evolution of magnetised, rotating stars and planets.

astro-ph.EP

Near-core magnetic field strengths inferred from gravity modes in intermediate-mass stars

In this work, we derive upper limits for the strength of the near-core magnetic field in intermediate-mass stars, since high-order g-modes can be fully suppressed by a critical magnetic field. Both poloidal and toroidal components of the magnetic field are included. We examine how the upper limits on magnetic field strengths are affected by the degree and azimuthal order of the oscillations, as well as the magnetic field configuration. We consider two gamma-Doradus stars hosting high-order g-modes and an evolved delta-Scuti star with mixed modes, all with prior mode identification from observations. We determine the best structural model from their stellar parameters through grid-based modeling with MESA. Frequencies for the best models are extracted using GYRE and matched to the observed modes. The critical magnetic fields for all calculated frequencies in our models are obtained from the Dedalus code, from which we can infer an upper limit on the near-core field strength. We find an upper limit on the near-core radial field strength of Br ~ 130 kG and Br ~ 13 kG, assuming a dipole field configuration, for the two gamma-Doradus stars KIC 3127996 and KIC 5876187, respectively. For 44 Tau, analysis of mixed modes yields a field strength of Br ~ 1771 kG. Different magnetic field configurations and mode degrees lead to different estimates. The results for the radial component of the magnetic field in the main sequence gamma-Doradus stars are consistent with estimates of magnetic field strengths in red giant stars that assume an internal field generated by a core dynamo, although the stronger of the two inferred magnetic fields may require some enhancement by a fossil field. The toroidal component does not affect g-modes significantly and is required to be more than 200 times stronger than the radial component to suppress g-modes. (abridged for arXiv)

astro-ph.SR

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy. To bridge this gap, we propose UniSinger, the first end-to-end framework unifying speaker cloning song generation and accompaniment co-generation SVC. Building on the multimodal diffusion transformer, we construct a unified speaker embedding space transferring speaker representation from SVC to song generation, endowing fine-grained cross-task timbre control. To mitigate multi-task optimization conflicts, we design a curriculum learning strategy using task-specific modality masking to guide the model to gradually master the generative mechanisms among semantic content, vocal timbre, and accompaniment. Experiments show state-of-the-art performance on both tasks and realizes complementary benefits, offering new possibilities for intelligent music production.

cs.SD

Inferring main-sequence stage and buoyancy-glitch amplitudes from Fourier spectra of gravity-mode period spacings: Ensemble Analysis of 26 Slowly Pulsating B Stars

Gravito-inertial-mode asteroseismology of intermediate-mass main-sequence stars took off with the 5-month uninterrupted light curves of the CoRoT space mission. It was developed in detail from the 4-year-long Kepler light curves, which provided a practical means to measure the rotation frequency in the transition layer between the convective core and the radiative envelope, where the local buoyancy frequency reaches a maximum. Recently, a new buoyancy glitch inversion method based on the Fourier spectra of gravity-mode period spacings was developed to probe that region further (Guo 2025). We aim to exploit the information contained in the variability of gravity-mode period spacings ($\Delta P$) in Slowly Pulsating B (SPB) stars with rotation. We investigate how well the main-sequence evolutionary stage can be inferred from this variability. We extract the frequency and amplitude of the variability in $\Delta P$ from the Fourier spectrum (FT). Both the period spacing $\Delta P$ and its periodic perturbations $\delta P$ (deviations from their asymptotic values) are used. The measured dominant frequency of $\Delta P$ allows us to infer the central hydrogen mass fraction, $X_c$, which is a main-sequence age indicator. The inferred $X_c$ values from $FT(\Delta P)$ mostly agree with previous results reported in the literature based on forward modelling of individual identified mode frequencies. We find that the buoyancy glitches $\delta N/N$ in SPB stars are generally less than $2\%$ in amplitude. Ensemble asteroseismic modeling of gravity-mode pulsators can now be carried out efficiently with our novel $FT(\Delta P)$ method once the internal rotation rate of the pulsators is known. Our methodology offers a fast method for gravito-inertial asteroseismic applications in the era of ongoing and future space-based observations.

astro-ph.SR

Asteroseismic Imprints of Mass Transfer in Binary Stars: Probing the Interiors of Donors and Accretors with Gravity and Acoustic Modes

Context. The synergy between close binary stars and asteroseismology enables constraints on mass-transfer episodes and their consequences for internal structure, rotation profiles, and oscillation modes. Aims. We investigate how mass accretion and donation in close binaries affects the internal structure and oscillation modes of main-sequence stars. Methods. Building on the established relation between the Brunt-Vaisala (buoyancy) glitch and the Fourier spectra of g-mode period spacings, we quantitatively explain the origins of the g-mode period-spacing differences between single-star and mass-accretion/donation models of intermediate-mass stars (M = 2.0, 3.0, and 4.5 Msun). In particular, the hydrogen mass fraction profiles X of the donor model show two chemical gradient regions, which results in a double-peaked Brunt-Vaisala profile. The presence of additional buoyancy glitches gives rise to further periodic modulations in the g-mode period spacings. Results. Mass-accretion induced changes in the chemical profile create sharp features in the buoyancy frequency, which modify both the amplitudes and frequencies of the g-mode period-spacing variations. This behavior resembles that produced by multiple chemical transition zones in compact pulsators such as white dwarfs and sub-dwarf B stars. Similarly, for acoustic modes in the M = 1 Msun solar-like models, we attribute the differences in frequency-separation ratios between single-star and mass-donor models to the variations in the internal sound-speed gradient (acoustic glitches). We discuss future prospects for using asteroseismology to discover the mass-transfer products and constrain the mass-transfer processes in binary star evolution.

astro-ph.SR

Asteroseismology and Buoyancy Glitch Inversion with Fourier Spectra of Gravity Mode Period Spacings

We investigate the small, quasi-periodic modulations seen in the gravity-mode period spacings of pulsating stars. These ``wiggles'' are produced by buoyancy glitches -- sharp features in the buoyancy frequency ($N$) caused by composition transitions and the convective-radiative interface. Our method takes the Fourier transform of the period-spacing series, $FT(\Delta P_k)$ as a function of radial order $k$. We show that $FT(\Delta P_k)$ traces the radial derivative of the normalized glitch profile $\delta N/N$ with respect to the normalized buoyancy radius; peaks in $FT(\Delta P_k)$ therefore pinpoint jump/drop locations in $N$ and measure their sharpness. We also note that the Fourier transform of relative period perturbations (deviations from asymptotic values), $FT(\delta P/P)$, directly recovers the absolute value of the glitch profile $|\delta N/N|$, enabling a straightforward inversion for the internal structure. The dominant $FT(\Delta P_k)$ frequency correlates tightly with central hydrogen abundance ($X_c$) and thus with stellar age for slowly pulsating B-stars, with only weak mass dependence. Applying the technique to MESA stellar models and to observed slowly pulsating B-stars and $\gamma$ Dor pulsators, we find typical glitch amplitudes $\delta N/N \lesssim 0.01$ and derivative magnitudes $\lesssim 0.1$, concentrated at chemical gradients and the convective boundary. This approach enables fast, ensemble asteroseismology of g-mode pulsators, constrains internal mixing and ages, and can be extended to other classes of pulsators, with potential links to tidal interactions in binaries.

astro-ph.SR

SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion

Voice Conversion (VC) aims to modify a speaker's timbre while preserving linguistic content. While recent VC models achieve strong performance, most struggle in real-time streaming scenarios due to high latency, dependence on ASR modules, or complex speaker disentanglement, which often results in timbre leakage or degraded naturalness. We present SynthVC, a streaming end-to-end VC framework that directly learns speaker timbre transformation from synthetic parallel data generated by a pre-trained zero-shot VC model. This design eliminates the need for explicit content-speaker separation or recognition modules. Built upon a neural audio codec architecture, SynthVC supports low-latency streaming inference with high output fidelity. Experimental results show that SynthVC outperforms baseline streaming VC systems in both naturalness and speaker similarity, achieving an end-to-end latency of just 77.1 ms.

cs.SD

WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction

In this paper, we present WEST(WE Speech Toolkit), a speech toolkit based on a large language model (LLM) for speech understanding, generation, and interaction. There are three key features of WEST: 1) Fully LLM-based: Standing on the shoulders of giants by reusing mature architectures, ecosystems (e.g., Hugging Face), and methods (e.g., sequence packing) from large models. 2) Full-stack: Supports tasks such as recognition, synthesis, understanding, dialogue, and multimodal capabilities, with extensibility to incorporate open-source models. 3) Simple and Stupid: A simple and stupid speech toolkit that everyone can Touch. In addition, WEST provides two types of recipes, models, and experimental results. The first is entirely based on open-source models and open-source data, allowing users to fully reproduce the experiments in this paper and serving as a verification system or minimal system baseline. The second is trained on massive data, offering superior performance so the user can directly apply it out of the box. WEST is publicly avilable at https://github.com/wenet-e2e/west/

cs.CL

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constructed using our novel Chuan-Pipeline, a complete data processing framework for dialectal speech. To facilitate rigorous evaluation and demonstrate the corpus's effectiveness, we also release high-quality ASR and TTS benchmarks, WenetSpeech-Chuan-Eval, with manually verified transcriptions. Experiments show that models trained on WenetSpeech-Chuan achieve state-of-the-art performance among open-source systems and demonstrate results comparable to commercial services. As the largest open-source corpus for Sichuanese dialects, WenetSpeech-Chuan not only lowers the barrier to research in dialectal speech processing but also plays a crucial role in promoting AI equity and mitigating bias in speech technologies. The corpus, benchmarks, models, and receipts are publicly available on our project page.

cs.CL

Asteroseismology of the young open cluster NGC 2516 II. Constraining cluster age using gravity-mode pulsators

Although asteroseismology is regarded as the most powerful tool for probing stellar interiors, seismic modelling remains dependent on global stellar parameters. Stellar clusters offer direct measurements of these parameters by fitting a CMD, making the application of asteroseismology in clusters a valuable approach to advancing stellar physics modelling. We aimed to develop seismic modelling for gravity-mode pulsators in the open cluster NGC 2516 to determine stellar ages. We computed 1D stellar models using MESA, incorporating rotation-induced transport processes. Exponential overshooting was included, as well as rotationally induced mixing in the radiative envelope. Grids of evolutionary models were computed covering isochrone-derived mass ranges. The models were evolved up to 300 Myr because of the cluster's young age (~100Myr). By fitting the frequencies of identified modes of four gravity-mode member pulsators simultaneously, we measure the seismic age of the cluster NGC 2516 as 132+-8Myr. This high-precision seismic age estimate deviates by 1sigma from the isochronal age derived from public MIST isochrones for rotating stars. Our findings show that seismic modelling strongly constrains core overshooting, but because the period spacing patterns are smooth, it provides weak constraints on mixing in the radiative envelopes. The two most massive gravity-mode pulsators have MIST masses ~2.0M_sun while their seismic masses are 1.75M_sun. We constructed new asteroseismology-calibrated isochrones using input physics identical to that of our seismic model grid. While this resolves the age discrepancy, the mass discrepancy is only partially addressed. The remaining small yet persisting mass discrepancy implies a mismatch between the physics in core to surface environments of 1D stellar models and the seismic observables probing those areas of fast-rotating stars.

astro-ph.SR

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. It comprises six modules: Audio Collection, Speaker Attributes Annotation, Speech Quality Annotation, Automatic Speech Recognition, Text Postprocessing and Recognizer Output Voting, enabling rich and high-quality annotations. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.

cs.SD

OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue

Empathy is crucial in enabling natural interactions within spoken dialogue systems, allowing machines to recognize and respond appropriately to paralinguistic cues such as age, gender, and emotion. Recent advancements in end-to-end speech language models, which unify speech understanding and generation, provide promising solutions. However, several challenges persist, including an over-reliance on large-scale dialogue datasets, insufficient extraction of paralinguistic cues vital for conveying empathy, and the lack of empathy-specific datasets and evaluation frameworks. To address these issues, we introduce OSUM-EChat, an open-source, end-to-end spoken dialogue system designed to enhance empathetic interactions, particularly in resource-limited settings. OSUM-EChat introduces two key innovations: (1) a three-stage understanding-driven spoken dialogue training strategy that extends the capabilities of a large speech understanding model to spoken dialogue tasks, and (2) a linguistic-paralinguistic dual thinking mechanism that integrates paralinguistic understanding through a chain of thought with dialogue generation, enabling the system to produce more empathetic responses. This approach reduces reliance on large-scale dialogue datasets while maintaining high-quality empathetic interactions. Additionally, we introduce the EChat-200K dataset, a rich corpus of empathetic speech-to-speech dialogues, and the EChat-eval benchmark, a comprehensive framework for evaluating the empathetic capabilities of dialogue systems. Experimental results demonstrate that OSUM-EChat outperforms end-to-end spoken dialogue models regarding empathetic responsiveness, validating its effectiveness.

cs.SD

The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge

This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a speech generator for Track 1 and a background audio generator for Track 2. In Track 1, we employ Single-Codec to tokenize the speech into discrete tokens and use a language-model-based approach to achieve zero-shot speaking style cloning. The Single-Codec effectively decouples timbre and speaking style at the token level, reducing the acoustic modeling burden on the autoregressive language model. Additionally, we use DSPGAN to upsample 16 kHz mel-spectrograms to high-fidelity 48 kHz waveforms. In Track 2, we propose a background audio generator based on large language models (LLMs). This system produces scene-appropriate accompaniment descriptions, synthesizes background audio with Tango 2, and integrates it with the speech generated by our Track 1 system. Our submission achieves the second place and the first place in Track 1 and Track 2 respectively.

cs.SD

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge: Tasks, Results and Findings

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneous behaviors in conversational speech. The challenge comprises two tracks: an unconstrained track without limitation on data and model usage, and a constrained track only allowing the use of constrained open-source datasets. A 100-hour high-quality conversational speech dataset is also made available with the challenge. This paper details the data, tracks, submitted systems, evaluation results, and findings.

cs.SD

Bird-inspired tendon coupling improves paddling efficiency by shortening phase transition times

Drag-based swimming with rowing appendages, fins, and webbed feet is a widely adapted locomotion form in aquatic animals. To develop effective underwater and swimming vehicles, a wide range of bioinspired drag-based paddles have been proposed, often faced with a trade-off between propulsive efficiency and versatility. Webbed feet provide an effective propulsive force in the power phase, are light weight and robust, and can even be partially folded away in the recovery phase. However, during the transition between recovery and power phase, much time is lost folding and unfolding, leading to drag and reducing efficiency. In this work, we took inspiration from the coupling tendons of aquatic birds and utilized tendon coupling mechanisms to shorten the transition time between recovery and power phase. Results from our hardware experiments show that the proposed mechanisms improve propulsive efficiency by 2.0 and 2.4 times compared to a design without extensor tendons or based on passive paddle, respectively. We further report that distal leg joint clutching, which has been shown to improve efficiency in terrestrial walking, did not play an major role in swimming locomotion. In sum, we describe a new principle for an efficient, drag-based leg and paddle design, with potential relevance for the swimming mechanics in aquatic birds.

cs.RO

Oscillation Frequencies of Moderately Rotating Delta Scuti Stars: Asymmetric Mode Splittings Due to Non-spherical Distortion

We find that the observed pressure-mode rotational splittings of slowly/moderately rotating Delta Scuti stars and Beta Cephei stars mostly have a positive asymmetry. That is, the left frequency spacing is larger than the right spacing in the dipole mode splitting triplets and the $l=2$ mode splitting multiplets (considering $m=1, 0, -1$ modes only). This is in agreement with the second-order perturbative effect of the rotational non-spherical distortion: both the prograde and retrograde modes have their frequencies shifted towards lower values relative to the $m=0$ modes. We thus study the rotational perturbation both in the first and second order, as well as the near-degeneracy mode coupling effect in MESA models representing Delta Scuti stars. For faster rotators, the near-degeneracy mode coupling between the nearest radial and quadrupole modes can significantly shift the $m=0$ modes, reduce the splitting asymmetry, and even change its sign. We find the theoretical splitting asymmetry from the second-order non-spherical distortion is larger than observed asymmetry. To facilitate future detections, we predict correlations between splitting asymmetry, splitting amplitude, and pulsation frequency. We also discuss additional factors that can influence splitting asymmetry, including embedded magnetic fields, resonant mode coupling, and binarity.

astro-ph.SR

Pulsation Phases and Mode Identification of Tidally Excited Oscillations in Fourteen Kepler Heartbeat Stars

Tidally excited oscillations (TEOs) in Heartbeat Stars (HBSs) are an essential probe of the internal properties of the systems, but their potential has yet to be fully exploited. Based on the orbital parameters of TEO candidates from our previous works, we identify the pulsation phases and amplitudes of TEOs in fourteen Kepler HBSs. Most pulsation phases of most systems can be explained by the dominant being $l=2$, $m=0$, or $\pm2$ spherical harmonic, assuming that the spin and orbital axes are aligned, and the pulsations are adiabatic and standing waves. The largest deviation ($>6\sigma$) occurs in KIC 8459354, which can be explained by the spin-orbit misalignment, and KIC 5877364 has a similar scenario. For KIC 11122789, almost half of the harmonics show large deviations; we cautiously suggest that these harmonics may not be considered TEO candidates. A similar scenario also exists in KIC 6290740. This phases and mode identification approach can also be used inversely to verify the TEO candidates derived by the Fourier analysis. Furthermore, the harmonics with large deviations ($>2\sigma$) in KIC 4377638, KIC 5090937, and KIC 11403032 can be expected to be travelling waves rather than standing waves. In addition, we also suggest that the apsidal motion could cause large deviations in TEO phases from theoretical values.

astro-ph.SR

Open-Structure: Structural Benchmark Dataset for SLAM Algorithms

This paper presents Open-Structure, a novel benchmark dataset for evaluating visual odometry and SLAM methods. Compared to existing public datasets that primarily offer raw images, Open-Structure provides direct access to point and line measurements, correspondences, structural associations, and co-visibility factor graphs, which can be fed to various stages of SLAM pipelines to mitigate the impact of data preprocessing modules in ablation experiments. The dataset comprises two distinct types of sequences from the perspective of scenarios. The first type maintains reasonable observation and occlusion relationships, as these critical elements are extracted from public image-based sequences using our dataset generator. In contrast, the second type consists of carefully designed simulation sequences that enhance dataset diversity by introducing a wide range of trajectories and observations. Furthermore, a baseline is proposed using our dataset to evaluate widely used modules, including camera pose tracking, parametrization, and factor graph optimization, within SLAM systems. By evaluating these state-of-the-art algorithms across different scenarios, we discern each module's strengths and weaknesses in the context of camera tracking and optimization processes. The Open-Structure dataset and baseline system are openly accessible on website: \url{https://open-structure.github.io}, encouraging further research and development in the field of SLAM.

cs.RO