SearcharxivSearch

arXiv subjects

Suntae Hwang

Publications and source records attributed to Suntae Hwang.

3 recordsLinked to original sources

DOSE : Drum One-Shot Extraction from Music Mixture

Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music mixture. To facilitate this, we propose the Random Mixture One-shot Dataset (RMOD), comprising large-scale, randomly arranged music mixtures paired with corresponding drum one-shot samples. Our proposed model, Drum One- Shot Extractor (DOSE), leverages neural audio codec language models for end-to-end extraction, bypassing traditional source separation steps. Additionally, we introduce a novel onset loss, designed to encourage accurate prediction of the initial transient of drum one-shots, which is essential for capturing timbral characteristics. We compare this approach against a source separation-based extraction method as a baseline. The results, evaluated using Frechet Audio Distance (FAD) and Multi-Scale Spectral loss (MSS), demonstrate that DOSE, enhanced with onset loss, outperforms the baseline, providing more accurate and higher-quality drum one-shots from music mixtures. The code, model checkpoint, and audio examples are available at https://github.com/HSUNEH/DOSE

eess.AS

Song Form-aware Full-Song Text-to-Lyrics Generation with Multi-Level Granularity Syllable Count Control

Lyrics generation presents unique challenges, particularly in achieving precise syllable control while adhering to song form structures such as verses and choruses. Conventional line-by-line approaches often lead to unnatural phrasing, underscoring the need for more granular syllable management. We propose a framework for lyrics generation that enables multi-level syllable control at the word, phrase, line, and paragraph levels, aware of song form. Our approach generates complete lyrics conditioned on input text and song form, ensuring alignment with specified syllable constraints. Generated lyrics samples are available at: https://tinyurl.com/lyrics9999

cs.CL

A Cyberinfrastructure-based Approach to Real Time Water Temperature Prediction

The prediction of water temperature is crucial for aquatic ecosystem studies and management. In this paper, we raise challenging issues in supporting real time water temperature prediction and present a system called WT-Agabus to address those issues. The WT-Agabus system is designed to be a cyberinfrastructure and to support various prediction models in a uniform way. In addition, we present a neural network-based water temperature prediction model to use only data available online from Korea Meteorological Administration (KMA). In this paper, we also show the current prototype implementation of the WT-Agabus system to support the prediction model

cs.DC