arXiv · 2110.12676
Controllable and Interpretable Singing Voice Decomposition via Assem-VC
Abstract
We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize the singing voice of the target speaker. In conclusion, we made a perfectly synced duet with the user's singing voice and the target singer's converted singing voice.
Explore related subjects
Keep this discovery
Kang-wook Kim, Junhyeok Lee. 2021-10-25. Controllable and Interpretable Singing Voice Decomposition via Assem-VC. https://arxiv.org/abs/2110.12676
Cite the original work for its findings. Save a collection to share your selection of sources.