SearcharxivSearch

arXiv subjects

Sunghyun Kim

Publications and source records attributed to Sunghyun Kim.

At least 19 recordsLinked to original sources

Microfabricated Au and Au/graphene bilayer platelets for levitation experiments

We describe a fabrication process for preparing liquid suspensions of micron-scale Au and Au/graphene bilayer platelets using thin-film deposition, optical lithography, ion milling, hydrofluoric acid (HF) substrate etching, and release from the substrate into a liquid suspension. Residual HF is removed through repeated centrifugation, decanting, and dilution cycles. The resulting suspension is characterized by electrospray deposition onto a secondary substrate, followed by electron and atomic force microscopy. The deposited platelets exhibit minimal aggregation, and the overall platelet yield reaches up to 30% of the platelets originally patterned on the wafer. Lateral force microscopy further confirms that the Au/graphene bilayer remains intact throughout fabrication, release, and electrospray deposition. This process provides a practical route for preparing high-quality platelet suspensions for levitated nanoparticle experiments and other applications requiring suspensions of two-dimensional nanostructures.

cond-mat.mes-hall

Carbon encapsulation of levitated Au nanoparticles

We investigate the formation of a barrier to evaporation that develops when levitated nanoscale Au nanoparticles are exposed to pulses of 532 nm laser radiation in a high vacuum (pressure $p=10^{-8}-10^{-7}$ Torr) environment. Our data are derived from precision measurements of the charge to mass ratio ($Q/M$) of $\sim$200 nm diameter Au particles confined in a quadrupole ion trap. We characterize the development of the barrier over time as the particle is repeatedly heated with laser pulses and determine the impact of variations of the interval between pulses and of exposure to several gases added to the vacuum chamber. We observe a slow increase in the mass of particles upon prolonged exposure to the vacuum, which we attribute to the growth of a barrier layer. For particles that have acquired a barrier during exposure to CO, we observe a rapid decrease in their mass upon subsequent exposure to O$_2$. These findings are consistent with the growth and subsequent oxidation of a graphene layer on the Au that forms the barrier to evaporation. However, we have not found that the rate of formation of the barrier depends on the pressure of carbon-containing gases (CO, C$_2$H$_4$, CO$_2$) we have added to the chamber. We hypothesize that a rare surface state on the solid Au particle catalyzes the reaction that introduces C to the particle. Repeated laser pulse heating is necessary--either to enable diffusion away from this state or to create fresh states that allow continued C uptake--to facilitate the growth of the surface graphene layer.

cond-mat.mes-hall

Collection, characterization, and precision measurement of levitated charged nanoparticles

We describe apparatus and experimental procedures for high stability precision measurements of levitated nanoscale particles confined in an ion trap in high vacuum. We discuss methods for particle generation and collection using electrospray emission, for rapid characterization by direct imaging of thermal motion, and for transfer of the particle from the trap where it is collected to a separate analysis trap in order to achieve better vacuum and lower noise. In the analysis trap at high vacuum (pressure $p\simeq10^{-8}$ Torr), we employ thermostatic control of the trapped particle oscillation amplitudes, allowing long-term, precision measurements of oscillation frequencies, from which the charge to mass ratio ($Q/M$) can be deduced. Under these conditions, we achieve $Q/M$ measurement precision approaching $10^{-5}$. This sensitivity will enable, for example, investigations of the surface chemistry of $\mu$m-scale levitated materials in ultra-high vacuum environments.

cond-mat.mes-hall

Equivariant Latent Alignment via Flow Matching under Group Symmetries

Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equivariant representation learning has emerged as a powerful framework for constructing latent spaces where analytically known group transformations could act directly, capturing geometric structure in data and enhancing both interpretability and generalization in novel view synthesis. However, we identify that existing approaches often suffer from latent misalignment, a discrepancy between the intended group action and the actually required transformations in the latent space. Consequently, the learned latents often fail to consistently preserve the equivariant relations imposed by the underlying group symmetry. To address this, we propose Residual Latent Flow, a flow-based framework that corrects the misaligned latents, thereby improving compliance with the underlying equivariance relation. Our comprehensive experiments show that our method significantly reduces latent misalignment and improves novel view synthesis quality, under rotation groups SO(n).

cs.CV

RecFlash: Fast Recommendation System on In-Storage Computing with Frequency-Based Data Mapping

Recommendation system has gained a large popularity for a variety of personalized suggestion tasks, but the ever-increasing number of user data makes real-time processing of recommendation systems difficult. NAND flash memory-based in-storage computing scheme can be one of favorable candidates among the various acceleration approaches because the flash memory typically has a larger memory capacity than the other memory types, so it can efficiently handle a large amount of user data for the recommendation inference services. However, different from other neural network applications where data is sequentially fetched from memory, the recommendation system shows the irregular random memory access pattern. Hence, most of the data loaded from the NAND flash array to the page buffer are not used, so a large portion of the internal bandwidth is underutilized, which degrades the performance on the inference acceleration of the recommendation tasks. In this paper, we propose RecFlash, a fast recommendation inference accelerator utilizing a data remapping algorithm with NAND flash-based in-storage computing (ISC). The experimental results show that our proposed method improves the latency and energy consumption by up to 81% and 91.9%, respectively, over the existing NAND flash-based ISC architecture.

cs.AR

Kinematic Modulation in Driven Spin Resonance

The transition probability of a spin driven by a rotating magnetic field is reformulated. This work shows that, once projection onto the measurement basis is properly accounted for, the laboratory measured probability is governed by both intrinsic spin dynamics and the time dependence of the measurement basis. For the rotating-field eigenbasis, this yields an additional kinematic modulation, leading to measurable deviations under strong driving. A unified probability expression is derived that subsumes the classic 1937 and 1954 formulations as limiting cases, while correcting the conventional treatment of magnetic resonance transitions.

quant-ph

DiSPA: Differential Substructure-Pathway Attention for Drug Response Prediction

Accurate prediction of drug response in precision medicine requires models that capture how specific chemical substructures interact with cellular pathway states. However, most existing deep learning approaches treat chemical and transcriptomic modalities independently or combine them only at late stages, limiting their ability to model fine-grained, context-dependent mechanisms of drug action. In addition, vanilla attention mechanisms are often sensitive to noise and sparsity in high-dimensional biological networks, hindering both generalization and interpretability. We present DiSPA (Differential Substructure-Pathway Attention), a framework that models bidirectional interactions between chemical substructures and pathway-level gene expression. DiSPA introduces differential cross-attention to suppress spurious associations while enhancing context-relevant interactions. On the GDSC benchmark, DiSPA achieves state-of-the-art performance, with strong improvements in the disjoint setting. These gains are consistent across random and drug-blind splits, suggesting improved robustness. Analyses of attention patterns indicate more selective and concentrated interactions compared to standard cross-attention. Exploratory evaluation shows that differential attention better prioritizes predefined target-related pathways, although this does not constitute mechanistic validation. DiSPA also shows promising generalization on external datasets (CTRP) and cross-dataset settings, although further validation is needed. It further enables zero-shot application to spatial transcriptomics, providing exploratory insights into region-specific drug sensitivity patterns without ground-truth validation.

cs.LG

MARBLE: Multi-Agent Reasoning for Bioinformatics Learning and Evolution

Motivation: Developing high-performing bioinformatics models typically requires repeated cycles of hypothesis formulation, architectural redesign, and empirical validation, making progress slow, labor-intensive, and difficult to reproduce. Although recent LLM-based assistants can automate isolated steps, they lack performance-grounded reasoning and stability-aware mechanisms required for reliable, iterative model improvement in bioinformatics workflows. Results: We introduce MARBLE, an execution-stable autonomous model refinement framework for bioinformatics models. MARBLE couples literature-aware reference selection with structured, debate-driven architectural reasoning among role-specialized agents, followed by autonomous execution, evaluation, and memory updates explicitly grounded in empirical performance. Across spatial transcriptomics domain segmentation, drug-target interaction prediction, and drug response prediction, MARBLE consistently achieves sustained performance improvements over strong baselines across multiple refinement cycles, while maintaining high execution robustness and low regression rates. Framework-level analyses demonstrate that structured debate, balanced evidence selection, and performance-grounded memory are critical for stable, repeatable model evolution, rather than single-run or brittle gains. Availability: Source code, data and Supplementary Information are available at https://github.com/PRISM-DGU/MARBLE.

cs.MA

GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism

Deep neural networks (DNNs) continue to grow rapidly in size, making them infeasible to train on a single device. Pipeline parallelism is commonly used in existing DNN systems to support large-scale DNN training by partitioning a DNN into multiple stages, which concurrently perform DNN training for different micro-batches in a pipeline fashion. However, existing pipeline-parallel approaches only consider sequential pipeline stages and thus ignore the topology of a DNN, resulting in missed model-parallel opportunities. This paper presents graph pipeline parallelism (GPP), a new pipeline-parallel scheme that partitions a DNN into pipeline stages whose dependencies are identified by a directed acyclic graph. GPP generalizes existing sequential pipeline parallelism and preserves the inherent topology of a DNN to enable concurrent execution of computationally-independent operators, resulting in reduced memory requirement and improved GPU performance. In addition, we develop GraphPipe, a distributed system that exploits GPP strategies to enable performant and scalable DNN training. GraphPipe partitions a DNN into a graph of stages, optimizes micro-batch schedules for these stages, and parallelizes DNN training using the discovered GPP strategies. Evaluation on a variety of DNNs shows that GraphPipe outperforms existing pipeline-parallel systems such as PipeDream and Piper by up to 1.6X. GraphPipe also reduces the search time by 9-21X compared to PipeDream and Piper.

cs.DC

Isometric Representation Learning for Disentangled Latent Space of Diffusion Models

The latent space of diffusion model mostly still remains unexplored, despite its great success and potential in the field of generative modeling. In fact, the latent space of existing diffusion models are entangled, with a distorted mapping from its latent space to image space. To tackle this problem, we present Isometric Diffusion, equipping a diffusion model with a geometric regularizer to guide the model to learn a geometrically sound latent space of the training data manifold. This approach allows diffusion models to learn a more disentangled latent space, which enables smoother interpolation, more accurate inversion, and more precise control over attributes directly in the latent space. Our extensive experiments consisting of image interpolations, image inversions, and linear editing show the effectiveness of our method.

cs.LG

An advance in the arithmetic of the Lie groups as an alternative to the forms of the Campbell-Baker-Hausdorff-Dynkin theorem

The exponential of an operator or matrix is widely used in quantum theory, but it sometimes can be a challenge to evaluate. For non-commutative operators ${\bf X}$ and ${\bf Y}$, according to the Campbell-Baker-Hausdorff-Dynkin theorem, ${\rm e}^{{\bf X}+{\bf Y}}$ is not equivalent to ${\rm e}^{\bf X}{\rm e}^{\bf Y}$, but is instead given by the well-known infinite series formula. For a Lie algebra of a basis of three operators $\{{\bf X,Y,Z}\}$, such that $[{\bf X}, {\bf Y}] = κ{\bf Z}$ for scalar $κ$ and cyclic permutations, here it is proven that ${\rm e}^{a{\bf X}+b{\bf Y}}$ is equivalent to ${\rm e}^{p{\bf Z}}{\rm e}^{q{\bf X}}{\rm e}^{-p{\bf Z}}$ for scalar $p$ and $q$. Extensions for ${\rm e}^{a{\bf X}+b{\bf Y}+c{\bf Z}}$ are also provided. This method is useful for the dynamics of atomic and molecular nuclear and electronic spins in constant and oscillatory transverse magnetic and electric fields.

quant-ph

Efficient Strong Scaling Through Burst Parallel Training

As emerging deep neural network (DNN) models continue to grow in size, using large GPU clusters to train DNNs is becoming an essential requirement to achieving acceptable training times. In this paper, we consider the case where future increases in cluster size will cause the global batch size that can be used to train models to reach a fundamental limit: beyond a certain point, larger global batch sizes cause sample efficiency to degrade, increasing overall time to accuracy. As a result, to achieve further improvements in training performance, we must instead consider "strong scaling" strategies that hold the global batch size constant and allocate smaller batches to each GPU. Unfortunately, this makes it significantly more difficult to use cluster resources efficiently. We present DeepPool, a system that addresses this efficiency challenge through two key ideas. First, burst parallelism allocates large numbers of GPUs to foreground jobs in bursts to exploit the unevenness in parallelism across layers. Second, GPU multiplexing prioritizes throughput for foreground training jobs, while packing in background training jobs to reclaim underutilized GPU resources, thereby improving cluster-wide utilization. Together, these two ideas enable DeepPool to deliver a 1.2 - 2.3x improvement in total cluster throughput over standard data parallelism with a single task when the cluster scale is large.

cs.DC

Nuclear Magnetic Resonance for Arbitrary Spin Values in the Rotating Wave Approximation

In order to probe the transitions of a nuclear spin $s$ from one of its substate quantum numbers $m$ to another substate $m'$, the experimenter applies a magnetic field ${\bm B}_0$ in some particular direction, such along $\hat{\bm z}$, and then applies an weaker field ${\bm B}_1(t)$ that is oscillatory in time with the angular frequency $ω$, and is normally perpendicular to ${\bm B}_0$, such as ${\bm B}_1(t)=B_1\hat{\bm x}\cos(ωt)$. In the rotating wave approximation, ${\bm B}_1(t)=B_1[\hat{\bm x}\cos(ωt)+\hat{\bm y}\sin(ωt)]$. Although this problem is solved for spin $\frac{1}{2}$ in every quantum mechanics textbook, for the general spin $s$ case, its general solution has been published only for the overall probability of a transition between the states, but the time dependence of the probability of finding the nucleus in each of the substates has not previously been published. Here we present an elementary method to solve this problem exactly, and present figures for the time dependencies of the various substates states for a variety of initial substate probabilities for a variety of $s$ values. We found a new result: unlike the $s=\frac{1}{2}$ case, for which if the initial probability of finding the particle in one of the substates was 1, and the time dependence of the probabilities of each of the substates oscillates between 0 and 1, for higher spin values, the time dependencies of the probabilities finding the particle in each of its substates, which periodic, is considerably more complicated.

quant-ph

Giant Huang-Rhys Factor for Electron Capture by the Iodine Interstitial in Perovskite Solar Cells

Improvement in the optoelectronic performance of halide perovskite semiconductors requires the identification and suppression of non-radiative carrier trapping processes. The iodine interstitial has been established as a deep level defect, and implicated as an active recombination centre. We analyse the quantum mechanics of carrier trapping. Fast and irreversible electron capture by the neutral iodine interstitial is found. The effective Huang-Rhys factor exceeds 300, indicative of the strong electron-phonon coupling that is possible in soft semiconductors. The accepting phonon mode has a frequency of 53 cm$^{-1}$ and has an associated electron capture coefficient of 10$^{-10}$cm$^3$s$^{-1}$. The inverse participation ratio is used to quantify the localisation of phonon modes associated with the transition. We infer that suppression of octahedral rotations is an important factor to enhance defect tolerance.

cond-mat.mtrl-sci

Ab initio calculation of the detailed balance limit to the photovoltaic efficiency of single p-n junction kesterite solar cells

The thermodynamic limit of photovoltaic efficiency for a single-junction solar cell can be readily predicted using the bandgap of the active light absorbing material. Such an approach overlooks the energy loss due to non-radiative electron-hole processes. We propose a practical ab initio procedure to determine the maximum efficiency of a thin-film solar cell that takes into account both radiative and non-radiative recombination. The required input includes the frequency-dependent optical absorption coefficient, as well as the capture cross-sections and equilibrium populations of point defects. For kesterite-structured Cu$_2$ZnSnS$_4$, the radiative limit is reached for a film thickness of around 2.6 micrometer, where the efficiency gain due to light absorption is counterbalanced by losses due to the increase in recombination current.

physics.app-ph

Assessing the defect tolerance of kesterite-inspired solar absorbers

Various thin-film I$_2$-II-IV-VI$_4$ photovoltaic absorbers derived from kesterite Cu$_2$ZnSn(S,Se)$_4$ have been synthesized, characterized, and theoretically investigated in the past few years. The availability of this homogeneous materials dataset is an opportunity to examine trends in their defect properties and identify criteria to find new defect-tolerant materials in this vast chemical space. We find that substitutions on the Zn site lead to a smooth decrease in band tailing as the ionic radius of the substituting cation increases. Unfortunately, this substitution strategy does not ensure the suppression of deeper defects and non-radiative recombination. Trends across the full dataset suggest that Gaussian and Urbach band tails in kesterite-inspired semiconductors are two separate phenomena caused by two different antisite defect types. Deep Urbach tails are correlated with the calculated band gap narrowing caused by the (2I$_\mathrm{II}$+IV$_\mathrm{II}$) defect cluster. Shallow Gaussian tails are correlated with the energy difference between the kesterite and stannite polymorphs, which points to the role of (I$_\mathrm{II}$+II$_\mathrm{I}$) defect clusters involving Group IB and Group IIB atoms swapping across \textit{different} cation planes. This finding can explain why \textit{in-plane} cation disorder and band tailing are uncorrelated in kesterites. Our results provide quantitative criteria for discovering new kesterite-inspired photovoltaic materials with low band tailing.

physics.app-ph

Quick-start guide for first-principles modelling of point defects in crystalline materials

Defects influence the properties and functionality of all crystalline materials. For instance, point defects participate in electronic (e.g. carrier generation and recombination) and optical (e.g. absorption and emission) processes critical to solar energy conversion. Solid-state diffusion, mediated by the transport of charged defects, is used for electrochemical energy storage. First-principles calculations of defects based on density functional theory have been widely used to complement, and even validate, experimental observations. In this `quick-start guide', we discuss the best practice in how to calculate the formation energy of point defects in crystalline materials and analysis techniques appropriate to probe changes in structure and properties relevant across energy technologies.

cond-mat.mtrl-sci

Comment on "Low-frequency lattice phonons in halide perovskites explain high defect tolerance toward electron-hole recombination"

Halide perovskites exhibit slow rates of non-radiative electron-hole recombination upon illumination. Chu et al. [Sci. Adv. 6 7, eaaw7453 (2020)] use the results of first-principles simulations to argue that this arises from the nature of the crystal vibrations and leads to a breakdown of Shockley-Read-Hall theory. We highlight flaws in their methodology and analysis of carrier capture by point defects in crystalline semiconductors.

cond-mat.mtrl-sci