SearcharxivSearch

arXiv subjects

Sooyoung Kim

Publications and source records attributed to Sooyoung Kim.

11 recordsLinked to original sources

Repurposing Image Diffusion Models for Training-Free Music Style Transfer on Mel-spectrograms

Music style transfer blends source structure with reference style to enable personalized music creation. However, existing zero-shot methods often struggle to capture fine-grained audio nuances, relying on coarse text descriptions or requiring expensive task-specific training. We propose Stylus, a training-free framework that repurposes pretrained image diffusion models for music style transfer in the Mel-spectrogram domain. By treating audio as structured time-frequency images, Stylus manipulates self-attention by injecting style keys and values while preserving source structural queries. To ensure high fidelity, we introduce a phase-preserving reconstruction strategy to mitigate spectrogram inversion artifacts, alongside a classifier-free-guidance-inspired control for adjustable stylization. Extensive evaluations including 2,925 human ratings demonstrate that Stylus outperforms state-of-the-art baselines, achieving 34.1% higher content preservation and 25.7% better perceptual quality. Our work validates that generic image priors can be effectively leveraged for the training-free transformation of structured Mel-spectrograms. Code and materials are available at https://github.com/Sooyyoungg/Stylus.git.

cs.SD

Macro2Micro: A Rapid and Precise Cross-modal Magnetic Resonance Imaging Synthesis using Multi-scale Structural Brain Similarity

The human brain is a complex system requiring both macroscopic and microscopic components for comprehensive understanding. However, mapping nonlinear relationships between these scales remains challenging due to technical limitations and the high cost of multimodal Magnetic Resonance Imaging (MRI) acquisition. To address this, we introduce Macro2Micro, a deep learning framework that predicts brain microstructure from macrostructure using a Generative Adversarial Network (GAN). Based on the hypothesis that microscale structural information can be inferred from macroscale structures, Macro2Micro explicitly encodes multiscale brain information into distinct processing branches. To enhance artifact elimination and output quality, we propose a simple yet effective auxiliary discriminator and learning objective. Extensive experiments demonstrated that Macro2Micro faithfully translates T1-weighted MRIs into corresponding Fractional Anisotropy (FA) images, achieving a 6.8\% improvement in the Structural Similarity Index Measure (SSIM) compared to previous methods, while retaining the individual biological characteristics of the brain. With an inference time of less than 0.01 seconds per MR modality translation, Macro2Micro introduces the potential for real-time multimodal rendering in medical and research applications. The code will be made available upon acceptance.

eess.IV

Revisiting Your Memory: Reconstruction of Affect-Contextualized Memory via EEG-guided Audiovisual Generation

In this paper, we introduce RevisitAffectiveMemory, a novel task designed to reconstruct autobiographical memories through audio-visual generation guided by affect extracted from electroencephalogram (EEG) signals. To support this pioneering task, we present the EEG-AffectiveMemory dataset, which encompasses textual descriptions, visuals, music, and EEG recordings collected during memory recall from nine participants. Furthermore, we propose RYM (Revisit Your Memory), a three-stage framework for generating synchronized audio-visual contents while maintaining dynamic personal memory affect trajectories. Experimental results demonstrate our method successfully decodes individual affect dynamics trajectories from neural signals during memory recall (F1=0.9). Also, our approach faithfully reconstructs affect-contextualized audio-visual memory across all subjects, both qualitatively and quantitatively, with participants reporting strong affective concordance between their recalled memories and the generated content. Especially, contents generated from subject-reported affect dynamics showed higher correlation with participants' reported affect dynamics trajectories (r=0.265, p<.05) and received stronger user preference (preference=56%) compared to those generated from randomly reordered affect dynamics. Our approaches advance affect decoding research and its practical applications in personalized media creation via neural-based affect comprehension. Codes and the dataset are available at https://github.com/ioahKwon/Revisiting-Your-Memory.

cs.AI

AesFA: An Aesthetic Feature-Aware Arbitrary Neural Style Transfer

Neural style transfer (NST) has evolved significantly in recent years. Yet, despite its rapid progress and advancement, existing NST methods either struggle to transfer aesthetic information from a style effectively or suffer from high computational costs and inefficiencies in feature disentanglement due to using pre-trained models. This work proposes a lightweight but effective model, AesFA -- Aesthetic Feature-Aware NST. The primary idea is to decompose the image via its frequencies to better disentangle aesthetic styles from the reference image while training the entire model in an end-to-end manner to exclude pre-trained models at inference completely. To improve the network's ability to extract more distinct representations and further enhance the stylization quality, this work introduces a new aesthetic feature: contrastive loss. Extensive experiments and ablations show the approach not only outperforms recent NST methods in terms of stylization quality, but it also achieves faster inference. Codes are available at https://github.com/Sooyyoungg/AesFA.

cs.CV

Energy-Efficient Downlink Semantic Generative Communication with Text-to-Image Generators

In this paper, we introduce a novel semantic generative communication (SGC) framework, where generative users leverage text-to-image (T2I) generators to create images locally from downloaded text prompts, while non-generative users directly download images from a base station (BS). Although generative users help reduce downlink transmission energy at the BS, they consume additional energy for image generation and for uploading their generator state information (GSI). We formulate the problem of minimizing the total energy consumption of the BS and the users, and devise a generative user selection algorithm. Simulation results corroborate that our proposed algorithm reduces total energy by up to 54% compared to a baseline with all non-generative users.

cs.LG

Integrating Voice-Based Machine Learning Technology into Complex Home Environments

To demonstrate the value of machine learning based smart health technologies, researchers have to deploy their solutions into complex real-world environments with real participants. This gives rise to many, oftentimes unexpected, challenges for creating technology in a lab environment that will work when deployed in real home environments. In other words, like more mature disciplines, we need solutions for what can be done at development time to increase success at deployment time. To illustrate an approach and solutions, we use an example of an ongoing project that is a pipeline of voice based machine learning solutions that detects the anger and verbal conflicts of the participants. For anonymity, we call it the XYZ system. XYZ is a smart health technology because by notifying the participants of their anger, it encourages the participants to better manage their emotions. This is important because being able to recognize one's emotions is the first step to better managing one's anger. XYZ was deployed in 6 homes for 4 months each and monitors the emotion of the caregiver of a dementia patient. In this paper we demonstrate some of the necessary steps to be accomplished during the development stage to increase deployment time success, and show where continued work is still necessary. Note that the complex environments arise both from the physical world and from complex human behavior.

cs.HC

Nonlinear Color-Metallicity Relations of Globular Clusters. X. Subaru/FOCAS Multi-object Spectroscopy of M87 Globular Clusters

We obtained spectra of some 140 globular clusters (GCs) associated with the Virgo central cD galaxy M87 with the Subaru/FOCAS MOS mode. The fundamental properties of GCs such as age, metallicity and $α$-element abundance are investigated by using simple stellar population models. It is confirmed that the majority of M87 GCs are as old as, more metal-rich than, and more enhanced in $α$-elements than the Milky Way GCs. Our high-quality, homogeneous dataset enables us to test the theoretical prediction of inflected color$-$metallicity relations (CMRs). The nonlinear-CMR hypothesis entails an alternative explanation for the widely observed GC color bimodality, in which even a unimodal metallicity spread yields a bimodal color distribution by virtue of nonlinear metallicity-to-color conversion. The newly derived CMRs of old, high-signal-to-noise-ratio GCs in M87 (the $V-I$ CMR of 83 GCs and the $M-T2$ CMR of 78 GCs) corroborate the presence of the significant inflection. Furthermore, from a combined catalog with the previous study on M87 GC spectroscopy, we find that a total of 185 old GCs exhibit a broad, unimodal metallicity distribution. The results corroborate the nonlinear-CMR interpretation of the GC color bimodality, shedding further light on theories of galaxy formation.

astro-ph.GA

Scalable Traffic Predictive Analysis using GPU in Big Data

The paper adopts parallel computing systems for predictive analysis in both CPU and GPU leveraging Spark Big Data platform. The traffic dataset is adopted to predict the traffic jams in Los Angeles County. It is collected from a popular platform in the USA for tracking information on the road using the device information and reports shared by the users. Large-scale traffic data set can be stored and processed using both GPU and CPU in this Scalable Big Data systems. The major contribution of this paper is to improve the performance of machine learning in distributed parallel computing systems with GPU to predict the traffic congestion. We show that the parallel computing can be achieve using both GPU and CPU with the existing Apache Spark platform. Our method can be applicable to other large scale datasets in different domains. The process modeling, as well as results, are interpreted using computing time and metrics: AUC, Precision and Recall. It should help the traffic management in Smart City.

cs.DC

Understanding Editing Behaviors in Multilingual Wikipedia

Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different languages. We take a computational approach to studying multilingualism using one of the largest user-generated content platforms, Wikipedia. We study multilingualism by collecting and analyzing a large dataset of the content written by multilingual editors of the English, German, and Spanish editions of Wikipedia. This dataset contains over two million paragraphs edited by over 15,000 multilingual users from July 8 to August 9, 2013. We analyze these multilingual editors in terms of their engagement, interests, and language proficiency in their primary and non-primary (secondary) languages and find that the English edition of Wikipedia displays different dynamics from the Spanish and German editions. Users primarily editing the Spanish and German editions make more complex edits than users who edit these editions as a second language. In contrast, users editing the English edition as a second language make edits that are just as complex as the edits by users who primarily edit the English edition. In this way, English serves a special role bringing together content written by multilinguals from many language editions. Nonetheless, language remains a formidable hurdle to the spread of content: we find evidence for a complexity barrier whereby editors are less likely to edit complex content in a second language. In addition, we find that multilinguals are less engaged and show lower levels of language proficiency in their second languages. We also examine the topical interests of multilingual editors and find that there is no significant difference between primary and non-primary editors in each language.

cs.SI

Nonlinear Color-Metallicity Relations of Globular Clusters. V. Nonlinear Absorption-line Index versus Metallicity Relations and Bimodal Index Distributions of M31 Globular Clusters

Recent spectroscopy on the globular cluster (GC) system of M31 with unprecedented precision witnessed a clear bimodality in absorption-line index distributions of old GCs. Such division of extragalactic GCs, so far asserted mainly by photometric color bimodality, has been viewed as the presence of merely two distinct metallicity subgroups within individual galaxies and forms a critical backbone of various galaxy formation theories. Given that spectroscopy is a more detailed probe into stellar population than photometry, the discovery of index bimodality may point to the very existence of dual GC populations. However, here we show that the observed spectroscopic dichotomy of M31 GCs emerges due to the nonlinear nature of metallicity-to-index conversion and thus one does not necessarily have to invoke two separate GC subsystems. We take this as a close analogy to the recent view that metallicity-color nonlinearity is primarily responsible for observed GC color bimodality. We also demonstrate that the metallicity-sensitive magnesium line displays non-negligible metallicity-index nonlinearity and Balmer lines show rather strong nonlinearity. This gives rise to bimodal index distributions, which are routinely interpreted as bimodal metallicity distributions, not considering metallicity-index nonlinearity. Our findings give a new insight into the constitution of M31's GC system, which could change much of the current thought on the formation of GC systems and their host galaxies.

astro-ph.GA

Nonlinear Color-Metallicity Relations of Globular Clusters. III. On the Discrepancy in Metallicity between Globular Cluster Systems and their Parent Elliptical Galaxies

One of the conundrums in extragalactic astronomy is the discrepancy in observed metallicity distribution functions (MDFs) between the two prime stellar components of early-type galaxies-globular clusters (GCs) and halo field stars. This is generally taken as evidence of highly decoupled evolutionary histories between GC systems and their parent galaxies. Here we show, however, that new developments in linking the observed GC colors to their intrinsic metallicities suggest nonlinear color-to-metallicity conversions, which translate observed color distributions into strongly-peaked, unimodal MDFs with broad metal-poor tails. Remarkably, the inferred GC MDFs are similar to the MDFs of resolved field stars in nearby elliptical galaxies and those produced by chemical evolution models of galaxies. The GC MDF shape, characterized by a sharp peak with a metal-poor tail, indicates a virtually continuous chemical enrichment with a relatively short timescale. The characteristic shape emerges across three orders of magnitude in the host galaxy mass, suggesting a universal process of chemical enrichment among various GC systems. Given that GCs are bluer than field stars within the same galaxy, it is plausible that the chemical enrichment processes of GCs ceased somewhat earlier than that of field stellar population, and if so, GCs preferentially trace the major, vigorous mode of star formation events in galactic formation. We further suggest a possible systematic age difference among GC systems, in that the GC systems in more luminous galaxies are older. This is consistent with the downsizing paradigm of galaxies and supports additionally the similar nature shared by GCs and field stars. Our findings suggest that GC systems and their parent galaxies have shared a more common origin than previously thought, and hence greatly simplify theories of galaxy formation.

astro-ph.CO