SearcharxivSearch

arXiv subjects

Sehwan Park

Publications and source records attributed to Sehwan Park.

5 recordsLinked to original sources

Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing

While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge. Existing pipelines predominantly yield flat, semantically agnostic collections of paths, where editing a single object requires manually identifying its constituent paths. To address this, we propose a VLM-driven agentic framework for semantic compositional SVG generation. Our pipeline recursively parses visual scenes into semantic and geometric hierarchies via top-down decomposition, visual grounding, and prompt-driven amodal occlusion recovery, ensuring each component is geometrically complete. Furthermore, we introduce the Semantic SVG Benchmark with human-annotated semantic groups and novel sub-component metrics (Semantic Recall/Precision, PERE) to explicitly evaluate structural compositionality and functional editability. Experiments show that our natively predicted structures surpass the upper bounds of existing flat-generation methods in both grouping quality and editability, while maintaining state-of-the-art visual fidelity.

cs.CV

Hot carrier diffusion-assisted ideal carrier multiplication in monolayer MoSe2

Carrier multiplication (CM), the process of generating multiple charge carriers from a single photon, offers an opportunity to exceed the Shockley-Queisser limit in photovoltaic applications. Despite extensive research, no material has yet achieved ideal CM efficiency, primarily owing to significant energy losses from carrier-lattice scattering. In this study, we demonstrate that monolayer MoSe2 can attain the theoretical maximum CM efficiency permitted by energy-momentum conservation principle, using ultrafast transient absorption spectroscopy. By resolving the scatter-free ballistic transport of hot carriers and validating our findings with first-principles calculations, we identify the cornerstone of optimal CM in monolayer MoSe2: superior hot-carrier dynamics characterized by suppressed energy dissipation via minimized carrier-lattice scattering, and the availability of abundant CM pathways facilitated by 2Eg band nesting. Comparative analysis with bulk MoSe2 further emphasizes the enhanced CM efficiency in the monolayer, attributed by superior hot-carrier diffusion and access to additional CM pathways. These results position monolayer MoSe2 as a promising candidate for high-performance optoelectronic applications, providing a robust platform for next-generation energy conversion technologies.

cond-mat.mes-hall

Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries

Recent advances in voice cloning and lip synchronization models have enabled Synthesized Audiovisual Forgeries (SAVFs), where both audio and visuals are manipulated to mimic a target speaker. This significantly increases the risk of misinformation by making fake content seem real. To address this issue, existing methods detect or localize manipulations but cannot recover the authentic audio that conveys the semantic content of the message. This limitation reduces their effectiveness in combating audiovisual misinformation. In this work, we introduce the task of Authentic Audio Recovery (AAR) and Tamper Localization in Audio (TLA) from SAVFs and propose a cross-modal watermarking framework to embed authentic audio into visuals before manipulation. This enables AAR, TLA, and a robust defense against misinformation. Extensive experiments demonstrate the strong performance of our method in AAR and TLA against various manipulations, including voice cloning and lip synchronization.

cs.SD

Random Conditioning with Distillation for Data-Efficient Diffusion Model Compression

Diffusion models generate high-quality images through progressive denoising but are computationally intensive due to large model sizes and repeated sampling. Knowledge distillation, which transfers knowledge from a complex teacher to a simpler student model, has been widely studied in recognition tasks, particularly for transferring concepts unseen during student training. However, its application to diffusion models remains underexplored, especially in enabling student models to generate concepts not covered by the training images. In this work, we propose Random Conditioning, a novel approach that pairs noised images with randomly selected text conditions to enable efficient, image-free knowledge distillation. By leveraging this technique, we show that the student can generate concepts unseen in the training images. When applied to conditional diffusion model distillation, our method allows the student to explore the condition space without generating condition-specific images, resulting in notable improvements in both generation quality and efficiency. This promotes resource-efficient deployment of generative diffusion models, broadening their accessibility for both research and real-world applications. Code, models, and datasets are available at https://dohyun-as.github.io/Random-Conditioning .

cs.LG

The low level RT control system of PLS-II storage ring at 400 mA 3.0 GeV

The RF system for the Pohang Light Source (PLS) storage ring was greatly upgraded for PLS-II project of 400mA, 3.0GeV from 200mA, 2.5GeV. Three superconducting(SC) RF cavities with each 300kW maximum klystron amplifier were commissioned with electron beam in way of one by one during the last 3 years for beam current of 400mA to until March 2014. The RF system is designed to provide stable beam through precise RF phase and amplitude requirements to be less than 0.3% in amplitude and 0.3° in phase deviations. This paper describes the RF system configuration, design details and test results.

physics.acc-ph