Searcharxiv⌕ Search

arXiv subjects

Yiping Ni

Publications and source records attributed to Yiping Ni.

2 recordsLinked to original sources

Temporal Anchors and Editing Sensitivity in Partial Speech Spoofing: A Controlled Study

Partial-spoof detectors must reject synthetic content while accepting benign edits. We diagnose temporal anchors and boundary-consistency training using frozen WavLM features, nonlinguistic unit controls, and source-locked evaluation. Genuine-genuine and genuine-fake splices contrast editing false alarms with synthetic-content misses; they do not isolate a unique causal artifact. With layer-6 features, phone units have higher PartialSpoof evaluation frame EER than soft frames (11.18% versus 9.59%). In 500 retrospective PS-eval cases, augmentation reduces genuine-splice false alarms; synthetic-core misses increase at 5% PS-development FPR but not demonstrably at 1%. A 200-case Llama A follow-up retains the false-alarm reduction but does not establish a miss-rate increase. Constructed-case ranking can improve while fixed-threshold misses rise. Consistency gives no uniform gain, and encoder-layer rankings change across corpora. We report frame and event localization with three seeds, including word units, and treat PS evaluation as retrospective. Audit code, selected training scripts, and result summaries are available at https://github.com/mysxs/partial-spoof-diagnostics.

cs.SD↗

VoiceWeaver: Staged Learning of Structured Controls for Expressive Speech and Sound-Event Generation

Adding expressive and environmental controls to a speech generator requires learning heterogeneous attributes without losing earlier capabilities. VoiceWeaver addresses this problem with structured label prefixes and staged emotion-tone-event training in a shared text-audio model. Replay-based distillation retains earlier predictions, while embedding decorrelation and attribute dropout regularize conditioning. Evaluation separates single-attribute correctness, joint emotion-event generation, and three-attribute correctness. Rounded to whole percentage points, emotion accuracy is 86% in Chinese and 83% in English. Chinese emotion-event joint accuracy is about 70%, versus 56% for mixed-task training and 48% for Ming-omni-tts; English joint accuracy is 66%. Three-attribute accuracy is 59% in Chinese and 55% in English on 500 samples each (57% pooled). Text error rates exceed those of the external TTS baselines. Audio samples are available at https://anonymous.4open.science/api/repo/VoiceWeaver1-4D88/file/index.html.

cs.SD↗