Searcharxiv⌕ Search

arXiv subjects

Sang Hyun Park

Publications and source records attributed to Sang Hyun Park.

At least 19 recordsLinked to original sources

Tunable Superconductivity Mediated by Heavy-Electron Plasmons: Band-Structure and Quantum-Geometric Engineering

Conventional superconductivity derives its pairing glue from lattice vibrations, tying its characteristic scales to chemistry and atomic masses. Plasmons$-$the collective oscillations of electrons$-$can instead be reshaped through electronic structure engineering, but the principles governing optimal plasmon-mediated pairing remain unclear. Here, we establish such principles for two-carrier systems in which heavy-electron plasmons mediate the pairing of light electrons. Within the random-phase approximation and Eliashberg theory, we calculate the optimal $T_c$ of minimal metallic models and show that it is controlled by a competition between the plasmon energy scale and retardation-driven suppression of the repulsion, yielding optimal carrier densities and band masses. While the plasmon channel alone reaches only $T_c\sim$ 0.1 K, a moderate phonon attraction cooperates with it, boosting $T_c$ by two orders of magnitude to above 20 K. However, the band flattening needed for slow metallic plasmons also favors the development of competing orders. We therefore consider an insulating system in which coherent interband transitions between flat bands generate gapped interband plasmons without free carriers. The heavy-band quantum metric governs the dispersion and electron-plasmon pairing strength of the interband plasmon, while the quantum geometry of the light band suppresses static screening and enhances the net attraction. Because layer separation rapidly weakens pairing, we propose systems with coexisting light and heavy electrons living in different mirror-symmetry sectors of the same layer as promising platforms. Our results establish a new role for flat-band systems in superconductivity: rather than hosting the paired electrons themselves, they can serve as a tunable pairing mediator whose collective charge excitations set the superconducting energy scale beyond their narrow bandwidth.

cond-mat.supr-con↗

Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following

Large vision language models (LVLMs) have demonstrated impressive performance across a wide range of tasks. These capabilities largely stem from visual instruction tuning, which fine-tunes models on datasets consisting of curated image-instruction-output triplets. However, in the medical domain, constructing large-scale, high-quality instruction datasets is particularly challenging due to the need for specialized expert knowledge. To address this issue, we propose an instruction-free tuning approach that reduces reliance on handcrafted instructions, leveraging only image-description pairs for fine-tuning. Specifically, we introduce a momentum proxy instruction as a replacement for curated text instructions, which preserves the instruction-following capability of the pre-trained LVLM while promoting updates to parameters that remain valid during inference. Consequently, the fine-tuned LVLM can flexibly respond to domain-specific instructions, even though explicit instructions are absent during fine-tuning. Additionally, we incorporate a response shuffling strategy to mitigate the model's over-reliance on previous words, facilitating more effective fine-tuning. Our approach achieves state-of-the-art accuracy on multiple-choice visual question answering tasks across SKINCON, WBCAtt, CBIS, and MIMIC-CXR datasets, significantly enhancing the fine-tuning efficiency of LVLMs in medical domains.

cs.CV↗

Enhanced enantiomer discrimination with chiral surface plasmons

Strong light-matter coupling in chiral cavities has been proposed as an effective way to selectively interact with an enantiomer that shares the same handedness as the cavity's chiral mode. We show that surface plasmons supported by a two-dimensional interface with both electric and chiral conductivities discriminate enantiomers more efficiently than chiral optical cavities. A quantum-electrodynamic treatment is developed to incorporate the molecule's electric and magnetic dipole moments. We show that the discrimination factor for a chiral plasmon can exceed that of the best chiral-mirror cavity by almost an order of magnitude due to stronger field confinement. In addition, surface plasmons couple to a dipole's projection onto an entire plane, whereas cavity (or free-space) modes couple only to a single polarization axis. This geometric difference produces a $\sqrt{2}$ orientation-averaged boost in chiral discrimination for chiral surface platforms. A handedness-preserving reflector further amplifies the enhancement, opening a practical route towards chiral sensing using twisted-layer platforms.

physics.optics↗

Model Agnostic Preference Optimization for Medical Image Segmentation

Preference optimization offers a scalable supervision paradigm based on relative preference signals, yet prior attempts in medical image segmentation remain model-specific and rely on low-diversity prediction sampling. In this paper, we propose MAPO (Model-Agnostic Preference Optimization), a training framework that utilizes Dropout-driven stochastic segmentation hypotheses to construct preference-consistent gradients without direct ground-truth supervision. MAPO is fully architecture- and dimensionality-agnostic, supporting 2D/3D CNN and Transformer-based segmentation pipelines. Comprehensive evaluations across diverse medical datasets reveal that MAPO consistently enhances boundary adherence, reduces overfitting, and yields more stable optimization dynamics compared to conventional supervised training.

cs.CV↗

Temporal Grounding as a Learning Signal for Referring Video Object Segmentation

Referring Video Object Segmentation (RVOS) aims to segment and track objects in videos based on natural language expressions, requiring precise alignment between visual content and textual queries. However, existing methods often suffer from semantic misalignment, largely due to indiscriminate frame sampling and supervision of all visible objects during training -- regardless of their actual relevance to the expression. We identify the core problem as the absence of an explicit temporal learning signal in conventional training paradigms. To address this, we introduce MeViS-M, a dataset built upon the challenging MeViS benchmark, where we manually annotate temporal spans when each object is referred to by the expression. These annotations provide a direct, semantically grounded supervision signal that was previously missing. To leverage this signal, we propose Temporally Grounded Learning (TGL), a novel learning framework that directly incorporates temporal grounding into the training process. Within this frame- work, we introduce two key strategies. First, Moment-guided Dual-path Propagation (MDP) improves both grounding and tracking by decoupling language-guided segmentation for relevant moments from language-agnostic propagation for others. Second, Object-level Selective Supervision (OSS) supervises only the objects temporally aligned with the expression in each training clip, thereby reducing semantic noise and reinforcing language-conditioned learning. Extensive experiments demonstrate that our TGL framework effectively leverages temporal signal to establish a new state-of-the-art on the challenging MeViS benchmark. We will make our code and the MeViS-M dataset publicly available.

cs.CV↗

JPEG Processing Neural Operator for Backward-Compatible Coding

Despite significant advances in learning-based lossy compression algorithms, standardizing codecs remains a critical challenge. In this paper, we present the JPEG Processing Neural Operator (JPNeO), a next-generation JPEG algorithm that maintains full backward compatibility with the current JPEG format. Our JPNeO improves chroma component preservation and enhances reconstruction fidelity compared to existing artifact removal methods by incorporating neural operators in both the encoding and decoding stages. JPNeO achieves practical benefits in terms of reduced memory usage and parameter count. We further validate our hypothesis about the existence of a space with high mutual information through empirical evidence. In summary, the JPNeO functions as a high-performance out-of-the-box image compression pipeline without changing source coding's protocol. Our source code is available at https://github.com/WooKyoungHan/JPNeO.

eess.IV↗

Latest Object Memory Management for Temporally Consistent Video Instance Segmentation

In this paper, we present Latest Object Memory Management (LOMM) for temporally consistent video instance segmentation that significantly improves long-term instance tracking. At the core of our method is Latest Object Memory (LOM), which robustly tracks and continuously updates the latest states of objects by explicitly modeling their presence in each frame. This enables consistent tracking and accurate identity management across frames, enhancing both performance and reliability through the VIS process. Moreover, we introduce Decoupled Object Association (DOA), a strategy that separately handles newly appearing and already existing objects. By leveraging our memory system, DOA accurately assigns object indices, improving matching accuracy and ensuring stable identity consistency, even in dynamic scenes where objects frequently appear and disappear. Extensive experiments and ablation studies demonstrate the superiority of our method over traditional approaches, setting a new benchmark in VIS. Notably, our LOMM achieves state-of-the-art AP score of 54.0 on YouTube-VIS 2022, a dataset known for its challenging long videos. Project page: https://seung-hun-lee.github.io/projects/LOMM/

cs.CV↗

Interfacial strong coupling and negative dispersion of propagating polaritons in freestanding oxide membranes

Membranes of complex oxides like perovskite SrTiO3 extend the multi-functional promise of oxide electronics into the nanoscale regime of two-dimensional materials. Here we demonstrate that free-standing oxide membranes supply a reconfigurable platform for nano-photonics based on propagating surface phonon polaritons. We apply infrared near-field imaging and -spectroscopy enabled by a tunable ultrafast laser to study pristine nano-thick SrTiO3 membranes prepared by hybrid molecular beam epitaxy. As predicted by coupled mode theory, we find that strong coupling of interfacial polaritons realizes symmetric and antisymmetric hybridized modes with simultaneously tunable negative and positive group velocities. By resolving reflection of these propagating modes from membrane edges, defects, and substrate structures, we quantify their dispersion with position-resolved nano-spectroscopy. Remarkably, we find polariton negative dispersion is both robust and tunable through choice of membrane dielectric environment and thickness and propose a novel design for in-plane Veselago lensing harnessing this control. Our work lays the foundation for tunable transformation optics at the nanoscale using polaritons in a wide range of freestanding complex oxide membranes.

cond-mat.mes-hall↗

Nodal lines in a honeycomb plasmonic crystal with synthetic spin

We analyze a plasmonic model on a honeycomb lattice of metallic nanodisks that hosts nodal lines protected by local symmetries. Using both continuum and tight-binding models, we show that a combination of a synthetic time-reversal symmetry, inversion symmetry, and particle-hole symmetry enforce the existence of nodal lines enclosing the $\mathrm{K}$ and $\mathrm{K}'$ points. The nodal lines are not directly gapped even when these symmetries are weakly broken. The existence of the nodal lines is verified using full-wave electromagnetic simulations. We also show that the degeneracies at nodal lines can be relieved by introducing a Kekulé distortion that acts to mix the nodal lines near the $\mathrm{K},\mathrm{K}'$ points. Our findings open pathways for designing novel plasmonic and photonic devices without reliance on complex symmetry engineering, presenting a convenient platform for studying nodal structures in two-dimensional systems.

physics.optics↗

CAT: Contrastive Adapter Training for Personalized Image Generation

The emergence of various adapters, including Low-Rank Adaptation (LoRA) applied from the field of natural language processing, has allowed diffusion models to personalize image generation at a low cost. However, due to the various challenges including limited datasets and shortage of regularization and computation resources, adapter training often results in unsatisfactory outcomes, leading to the corruption of the backbone model's prior knowledge. One of the well known phenomena is the loss of diversity in object generation, especially within the same class which leads to generating almost identical objects with minor variations. This poses challenges in generation capabilities. To solve this issue, we present Contrastive Adapter Training (CAT), a simple yet effective strategy to enhance adapter training through the application of CAT loss. Our approach facilitates the preservation of the base model's original knowledge when the model initiates adapters. Furthermore, we introduce the Knowledge Preservation Score (KPS) to evaluate CAT's ability to keep the former information. We qualitatively and quantitatively compare CAT's improvement. Finally, we mention the possibility of CAT in the aspects of multi-concept adapter and optimization.

cs.CV↗

Improving Text Generation on Images with Synthetic Captions

The recent emergence of latent diffusion models such as SDXL and SD 1.5 has shown significant capability in generating highly detailed and realistic images. Despite their remarkable ability to produce images, generating accurate text within images still remains a challenging task. In this paper, we examine the validity of fine-tuning approaches in generating legible text within the image. We propose a low-cost approach by leveraging SDXL without any time-consuming training on large-scale datasets. The proposed strategy employs a fine-tuning technique that examines the effects of data refinement levels and synthetic captions. Moreover, our results demonstrate how our small scale fine-tuning approach can improve the accuracy of text generation in different scenarios without the need of additional multimodal encoders. Our experiments show that with the addition of random letters to our raw dataset, our model's performance improves in producing well-formed visual text.

cs.CV↗

Illustrious: an Open Advanced Illustration Model

In this work, we share the insights for achieving state-of-the-art quality in our text-to-image anime image generative model, called Illustrious. To achieve high resolution, dynamic color range images, and high restoration ability, we focus on three critical approaches for model improvement. First, we delve into the significance of the batch size and dropout control, which enables faster learning of controllable token based concept activations. Second, we increase the training resolution of images, affecting the accurate depiction of character anatomy in much higher resolution, extending its generation capability over 20MP with proper methods. Finally, we propose the refined multi-level captions, covering all tags and various natural language captions as a critical factor for model development. Through extensive analysis and experiments, Illustrious demonstrates state-of-the-art performance in terms of animation style, outperforming widely-used models in illustration domains, propelling easier customization and personalization with nature of open source. We plan to publicly release updated Illustrious model series sequentially as well as sustainable plans for improvements.

cs.CV↗

Subject-Adaptive Transfer Learning Using Resting State EEG Signals for Cross-Subject EEG Motor Imagery Classification

Electroencephalography (EEG) motor imagery (MI) classification is a fundamental, yet challenging task due to the variation of signals between individuals i.e., inter-subject variability. Previous approaches try to mitigate this using task-specific (TS) EEG signals from the target subject in training. However, recording TS EEG signals requires time and limits its applicability in various fields. In contrast, resting state (RS) EEG signals are a viable alternative due to ease of acquisition with rich subject information. In this paper, we propose a novel subject-adaptive transfer learning strategy that utilizes RS EEG signals to adapt models on unseen subject data. Specifically, we disentangle extracted features into task- and subject-dependent features and use them to calibrate RS EEG signals for obtaining task information while preserving subject characteristics. The calibrated signals are then used to adapt the model to the target subject, enabling the model to simulate processing TS EEG signals of the target subject. The proposed method achieves state-of-the-art accuracy on three public benchmarks, demonstrating the effectiveness of our method in cross-subject EEG MI classification. Our findings highlight the potential of leveraging RS EEG signals to advance practical brain-computer interface systems. The code is available at https://github.com/SionAn/MICCAI2024-ResTL.

eess.SP↗

Few Shot Part Segmentation Reveals Compositional Logic for Industrial Anomaly Detection

Logical anomalies (LA) refer to data violating underlying logical constraints e.g., the quantity, arrangement, or composition of components within an image. Detecting accurately such anomalies requires models to reason about various component types through segmentation. However, curation of pixel-level annotations for semantic segmentation is both time-consuming and expensive. Although there are some prior few-shot or unsupervised co-part segmentation algorithms, they often fail on images with industrial object. These images have components with similar textures and shapes, and a precise differentiation proves challenging. In this study, we introduce a novel component segmentation model for LA detection that leverages a few labeled samples and unlabeled images sharing logical constraints. To ensure consistent segmentation across unlabeled images, we employ a histogram matching loss in conjunction with an entropy loss. As segmentation predictions play a crucial role, we propose to enhance both local and global sample validity detection by capturing key aspects from visual semantics via three memory banks: class histograms, component composition embeddings and patch-level representations. For effective LA detection, we propose an adaptive scaling strategy to standardize anomaly scores from different memory banks in inference. Extensive experiments on the public benchmark MVTec LOCO AD reveal our method achieves 98.1% AUROC in LA detection vs. 89.6% from competing methods.

cs.CV↗

Generating Realistic Brain MRIs via a Conditional Diffusion Probabilistic Model

As acquiring MRIs is expensive, neuroscience studies struggle to attain a sufficient number of them for properly training deep learning models. This challenge could be reduced by MRI synthesis, for which Generative Adversarial Networks (GANs) are popular. GANs, however, are commonly unstable and struggle with creating diverse and high-quality data. A more stable alternative is Diffusion Probabilistic Models (DPMs) with a fine-grained training strategy. To overcome their need for extensive computational resources, we propose a conditional DPM (cDPM) with a memory-efficient process that generates realistic-looking brain MRIs. To this end, we train a 2D cDPM to generate an MRI subvolume conditioned on another subset of slices from the same MRI. By generating slices using arbitrary combinations between condition and target slices, the model only requires limited computational resources to learn interdependencies between slices even if they are spatially far apart. After having learned these dependencies via an attention network, a new anatomy-consistent 3D brain MRI is generated by repeatedly applying the cDPM. Our experiments demonstrate that our method can generate high-quality 3D MRIs that share a similar distribution to real MRIs while still diversifying the training set. The code is available at https://github.com/xiaoiker/mask3DMRI_diffusion and also will be released as part of MONAI, at https://github.com/Project-MONAI/GenerativeModels.

eess.IV↗

Helical boundary modes from synthetic spin in a plasmonic lattice

Artificial lattices have been used as a platform to extend the application of topological physics beyond electronic systems. Here, using the two-dimensional Lieb lattice as a prototypical example, we show that an array of disks which each support localized plasmon modes give rise to an analog of the quantum spin Hall state enforced by a synthetic time reversal symmetry. We find that an effective next-nearest-neighbor coupling mechanism intrinsic to the plasmonic disk array introduces a nontrivial $Z_2$ topological order and gaps out the Bloch spectrum. A faithful mapping of the plasmonic system onto a tight-binding model is developed and shown to capture its essential topological signatures. Full wave numerical simulations of graphene disks arranged in a Lieb lattice confirm the existence of propagating helical boundary modes in the nontrivial band gap.

cond-mat.mes-hall↗

Content Preserving Image Translation with Texture Co-occurrence and Spatial Self-Similarity for Texture Debiasing and Domain Adaptation

Models trained on datasets with texture bias usually perform poorly on out-of-distribution samples since biased representations are embedded into the model. Recently, various image translation and debiasing methods have attempted to disentangle texture biased representations for downstream tasks, but accurately discarding biased features without altering other relevant information is still challenging. In this paper, we propose a novel framework that leverages image translation to generate additional training images using the content of a source image and the texture of a target image with a different bias property to explicitly mitigate texture bias when training a model on a target task. Our model ensures texture similarity between the target and generated images via a texture co-occurrence loss while preserving content details from source images with a spatial self-similarity loss. Both the generated and original training images are combined to train improved classification or segmentation models robust to inconsistent texture bias. Evaluation on five classification- and two segmentation-datasets with known texture biases demonstrates the utility of our method, and reports significant improvements over recent state-of-the-art methods in all cases.

cs.CV↗

Non-Hermitian chiral degeneracy of gated graphene metasurfaces

Non-Hermitian degeneracies, also known as exceptional points (EPs), have been the focus of much attention due to their singular eigenvalue surface structure. Nevertheless, as pertaining to a non-Hermitian metasurface platform, the reduction of an eigenspace dimensionality at the EP has been investigated mostly in a passive repetitive manner. Here, we propose an electrical and spectral way of resolving chiral EPs and clarifying the consequences of chiral mode collapsing of a non-Hermitian gated graphene metasurface. More specifically, the measured non-Hermitian Jones matrix in parameter space enables the quantification of nonorthogonality of polarisation eigenstates and half-integer topological charges associated with a chiral EP. Interestingly, the output polarisation state can be made orthogonal to the coalesced polarisation eigenstate of the metasurface, revealing the missing dimension at the chiral EP. In addition, the maximal nonorthogonality at the chiral EP leads to a blocking of one of the cross-polarised transmission pathways and, consequently, the observation of enhanced asymmetric polarisation conversion. We anticipate that electrically controllable non-Hermitian metasurface platforms can serve as an interesting framework for the investigation of rich non-Hermitian polarisation dynamics around chiral EPs.

physics.optics↗