SearcharxivSearch

arXiv subjects

Wenxin Fu

Publications and source records attributed to Wenxin Fu.

3 recordsLinked to original sources

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.

cs.SD

CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

Reference-conditioned melody-preserving lyric editing replaces words while retaining a performance's timing, singer identity, and naturalness. Continuous-latent autoregression avoids finite codebooks and offers stepwise generation with learned stopping. Editing creates a conflict absent from ordinary reconstruction: training pairs reference cues with original lyrics, whereas inference asks revised lyrics to override source-lyric-correlated cues; one source-following patch can propagate through AR history. We introduce CLASVS. Its State-Control-Transition (SCT) routing keeps target-lyric and reference-melody controls persistent, returns semantic feedback on phonetic progress to the causal planner, and confines the previous latent patch to the local Transition. Progressive State-Control Grounding (PSCG) learns this contract through paired-edit-free, content-consistent Mandarin reconstruction. On two Mandarin benchmarks, CLASVS improves all four operations over discrete-AR Vevo2 and reduces macro-PER by 46.2%, while maintaining melody, singer similarity, and perceptual quality. Together, these results establish a strong continuous-AR operating point for score-annotation-free lyric edits and a basis for broader stepwise control. Audio demonstrations are available on our project page: https://piedpiperg.github.io/clasvs-demo/.

cs.SD

On the maximal displacement of critical branching random walk in random environment

In this article, we study the maximal displacement of critical branching random walk in random environment. Let $M_n$ be the maximal displacement of a particle in generation $n$, and $Z_n$ be the total population in generation $n$, $M$ be the rightmost point ever reached by the branching random walk. Under some reasonable conditions, we prove a conditional limit theorem, \begin{equation*} \mathcal{L}\left( \dfrac{M_n}{\sqrt{\sigma} n^{\frac{3}{4}}} |Z_n>0\right) \dcon \mathcal{L}\left(A_\Lambda\right), \end{equation*} where random variable $A_\Lambda$ is related to the standard Brownian meander. And there exist some positive constant $C_1$ and $C_2$, such that \begin{equation*} C_1\leqslant\liminf\limits_{x\rightarrow\infty}x^{\frac{2}{3}}\P(M>x) \leqslant \limsup\limits_{x\rightarrow\infty} x^{\frac{2}{3}}\P(M>x) \leqslant C_2. \end{equation*} Compared with the constant environment case (Lalley and Shao (2015)), it revaels that, the conditional limit speed for $M_n$ in random environment (i.e., $n^{\frac{3}{4}}$) is significantly greater than that of constant environment case (i.e., $n^{\frac{1}{2}}$), and so is the tail probability for the $M$ (i.e., $x^{-\frac{2}{3}}$ vs $x^{-2}$). Our method is based on the path large deviation for the reduced critical branching random walk in random environment.

math.PR