SearcharxivSearch

arXiv subjects

Hyunseok Oh

Publications and source records attributed to Hyunseok Oh.

16 recordsLinked to original sources

Fine-tuning a vision-language model for fracture-surface morphology recognition

Vision-language models (VLMs) have shown strong potential for scientific image understanding, but general-purpose models often lack the domain-specific visual knowledge required for reliable materials characterization. In this work, we fine-tuned an open-source VLM (Qwen3-VL-32B-Instruct) for fracture-surface image analysis using a curated dataset of 13,168 open-source, literature-mined fracture-surface images. Morphology annotations were generated by GPT-5.2-Reasoning (high) from both the images and relevant excerpts of their source papers, and the dataset was further enriched with targeted manual collection and rotation-based augmentation. The resulting specialist model outperforms flagship proprietary multimodal models on a benchmark of 100 manually annotated images. It achieves a precision of 0.92, compared to 0.35 for the base Qwen3-VL-32B-Instruct, 0.58 for GPT-5.5-Reasoning (high), and 0.78 for Gemini 3.1 Pro-Reasoning (high). Dataset ablations show that manual collection of rare-feature images and augmentation via image rotation are both beneficial to improve recognition of less common fracture morphology features. We further discuss integrated use of the fine-tuned model with proprietary models to combine fracture-specific visual accuracy with broader multimodal reasoning for autonomous fractography. Although focused on fracture-surface images, this work demonstrates how VLMs can be adapted through targeted collection and fine-tuning on novel feature images to recognize those features and support downstream decision-making in autonomous microscopy workflows.

cond-mat.mtrl-sci

Neurosymbolic Language Reasoning as Satisfiability Modulo Theory

Natural language understanding requires interleaving textual and logical reasoning, yet large language models often fail to perform such reasoning reliably. Existing neurosymbolic systems combine LLMs with solvers but remain limited to fully formalizable tasks such as math or program synthesis, leaving natural documents with only partial logical structure unaddressed. We introduce Logitext, a neurosymbolic language that represents documents as natural language text constraints (NLTCs), making partial logical structure explicit. We develop an algorithm that integrates LLM-based constraint evaluation with satisfiability modulo theory (SMT) solving, enabling joint textual-logical reasoning. Experiments on a new content moderation benchmark, together with LegalBench and Super-Natural Instructions, show that Logitext improves both accuracy and coverage. This work is the first that treats LLM-based reasoning as an SMT theory, extending neurosymbolic methods beyond fully formalizable domains.

cs.AI

Quantum Nanophotonic Interface for Tin-Vacancy Centers in Thin-Film Diamond

The negatively charged tin-vacancy center in diamond (SnV$^-$) is an excellent solid state qubit with optically-addressable transitions and a long electron spin coherence time at elevated ($\sim1.7$ K). However, implementing scalable quantum nodes with high-fidelity optical readout of the electron spin state requires efficient photon emission and collection from the system. In this manuscript, we report a quantum photonic interface for SnV$^-$ centers based on one-dimensional photonic crystal cavities fabricated in diamond thin films. Furthermore, we provide a rigorous description of the spontaneous emission dynamics of our system, taking into account individual contributions from both the C and D transitions of the emitter. This allows for determination of Purcell factors per transition and, by extension, the C/D branching ratio SnV$^{-}$ zero phonon line. We observe quality factors up to $\sim$6000 across this sample, and measure up to a 12-fold lifetime reduction, which translates into a Purcell factor of $F_C=26.2\pm1.5$ for a targeted C transition. By considering the cavity mode polarization alignment with the C and D transition dipole moments, we validate the C/D branching ratio to be $\eta_{\text{BR}}=0.75\pm0.01$, in line with previous theoretical and experimental findings.

quant-ph

A spin-embedded diamond optomechanical resonator with mechanical quality factor exceeding one million

Diamond optomechanical crystal (OMC) devices with embedded color center spins are promising platforms for a broad range of applications in quantum sensing, networking, and computing applications, offering an interface between a GHz-frequency mechanical mode and both optical photons and coherent spins. A crucial but elusive step towards realizing this platform is to engineer a device with a high-quality factor mechanical mode while preserving the bulk-like coherence of embedded spins. Here we demonstrate sideband-resolved diamond OMCs with mechanical quality factors in excess of $10^6$ at cryogenic temperatures, and find coherence times up to $T_2$ = 270 $\mu$s for embedded nitrogen vacancy (NV) centers. Furthermore, we measure these devices across five orders of magnitude in intracavity optical power, demonstrating robust power handling and a high optomechanical cooperativity ($C\gg1$) at cryogenic temperatures that is essential for a broad range of quantum protocols requiring strong, coherent interactions between photons and phonons. These results are enabled by a robust, high-throughput method for forming single-crystal diamond membranes in combination with chemical vapor deposition (CVD) diamond overgrowth with nitrogen $\delta$-doping. We discuss the prospects of this platform for hybrid spin-mechanical devices in the quantum regime.

quant-ph

Thermal transport and the impact of hydrogen adsorption in Linde Type A zeolitic imidazolate frameworks

Thermal transport in metal-organic frameworks (MOFs) is of practical interest in diverse applications such as gas storage and separations, since insufficient heat dissipation can lead to detrimental effects. Despite investigations, influence of molecular infiltration on the heat transport remains unclear in many of MOFs due to poor understanding of mechanisms governing heat conductions. Here, we report molecular dynamics investigations of thermal transport properties in zeolitic imidazolate frameworks (ZIFs). We investigated Linde Type A topological ZIFs (ZIF-lta) exhibiting exceptionally low thermal conductivity with unusual trend of temperature dependence deviating from many crystalline materials, despite long-range crystalline order in them. We demonstrate that heat is predominantly carried by phonons with mean free paths comparable to their wavelengths, analogous to diffusons in amorphous solids owing to strong anharmonicity caused by complexity of unit cell consisting of a large number of metal centers. We further show that adsorbed hydrogen molecules increase thermal conductivity of ZIFs, mainly contributed by additional vibrational modes, as a result of gas-gas or gas-framework interactions. Our work advances fundamental understanding into the thermal transport in MOFs and suggests a means to engineer heat conduction via gas infiltrations.

cond-mat.mtrl-sci

Integrated phononic waveguide on thin-film lithium niobate on diamond

We demonstrate wavelength-scale phononic waveguides formed by transfer-printed thin-film lithium niobate (LN) on bulk diamond (LNOD), a material stack that combines the strong piezoelectricity of LN with the high acoustic velocity and color-center compatibility of diamond. We characterize a delay line based on a 100 micron long phononic waveguide at room and cryogenic temperatures. The total insertion loss through the device at 4 kelvin is -5.8 dB, corresponding to a >50% transducer efficiency, at a frequency of 2.8 gigahertz. Our work represents a step towards phonon-mediated hybrid quantum systems consisting of strain-sensitive color centers in diamond.

cond-mat.mes-hall

Beyond designer's knowledge: Generating materials design hypotheses via large language models

Materials design often relies on human-generated hypotheses, a process inherently limited by cognitive constraints such as knowledge gaps and limited ability to integrate and extract knowledge implications, particularly when multidisciplinary expertise is required. This work demonstrates that large language models (LLMs), coupled with prompt engineering, can effectively generate non-trivial materials hypotheses by integrating scientific principles from diverse sources without explicit design guidance by human experts. These include design ideas for high-entropy alloys with superior cryogenic properties and halide solid electrolytes with enhanced ionic conductivity and formability. These design ideas have been experimentally validated in high-impact publications in 2023 not available in the LLM training data, demonstrating the LLM's ability to generate highly valuable and realizable innovative ideas not established in the literature. Our approach primarily leverages materials system charts encoding processing-structure-property relationships, enabling more effective data integration by condensing key information from numerous papers, and evaluation and categorization of numerous hypotheses for human cognition, both through the LLM. This LLM-driven approach opens the door to new avenues of artificial intelligence-driven materials discovery by accelerating design, democratizing innovation, and expanding capabilities beyond the designer's direct knowledge.

cs.LG

Papez: Resource-Efficient Speech Separation with Auditory Working Memory

Transformer-based models recently reached state-of-the-art single-channel speech separation accuracy; However, their extreme computational load makes it difficult to deploy them in resource-constrained mobile or IoT devices. We thus present Papez, a lightweight and computation-efficient single-channel speech separation model. Papez is based on three key techniques. We first replace the inter-chunk Transformer with small-sized auditory working memory. Second, we adaptively prune the input tokens that do not need further processing. Finally, we reduce the number of parameters through the recurrent transformer. Our extensive evaluation shows that Papez achieves the best resource and accuracy tradeoffs with a large margin. We publicly share our source code at \texttt{https://github.com/snuhcs/Papez}

cs.SD

Sign Gradient Descent-based Neuronal Dynamics: ANN-to-SNN Conversion Beyond ReLU Network

Spiking neural network (SNN) is studied in multidisciplinary domains to (i) enable order-of-magnitudes energy-efficient AI inference and (ii) computationally simulate neuro-scientific mechanisms. The lack of discrete theory obstructs the practical application of SNN by limiting its performance and nonlinearity support. We present a new optimization-theoretic perspective of the discrete dynamics of spiking neurons. We prove that a discrete dynamical system of simple integrate-and-fire models approximates the sub-gradient method over unconstrained optimization problems. We practically extend our theory to introduce a novel sign gradient descent (signGD)-based neuronal dynamics that can (i) approximate diverse nonlinearities beyond ReLU and (ii) advance ANN-to-SNN conversion performance in low time steps. Experiments on large-scale datasets show that our technique achieves (i) state-of-the-art performance in ANN-to-SNN conversion and (ii) is the first to convert new DNN architectures, e.g., ConvNext, MLP-Mixer, and ResMLP. We publicly share our source code at https://github.com/snuhcs/snn_signgd .

cs.NE

Single-Spin Readout and Quantum Sensing using Optomechanically Induced Transparency

Solid-state spin defects are promising quantum sensors for a large variety of sensing targets. Some of these defects couple appreciably to strain in the host material. We propose to use this strain coupling for mechanically-mediated dispersive single-shot spin readout by an optomechanically-induced transparency measurement. Surprisingly, the estimated measurement times for negatively-charged silicon-vacancy defects in diamond are an order of magnitude shorter than those for single-shot optical fluorescence readout. Our scheme can also be used for general parameter-estimation metrology and offers a higher sensitivity than conventional schemes using continuous position detection.

quant-ph

On the Importance of Critical Period in Multi-stage Reinforcement Learning

The initial years of an infant's life are known as the critical period, during which the overall development of learning performance is significantly impacted due to neural plasticity. In recent studies, an AI agent, with a deep neural network mimicking mechanisms of actual neurons, exhibited a learning period similar to human's critical period. Especially during this initial period, the appropriate stimuli play a vital role in developing learning ability. However, transforming human cognitive bias into an appropriate shaping reward is quite challenging, and prior works on critical period do not focus on finding the appropriate stimulus. To take a step further, we propose multi-stage reinforcement learning to emphasize finding ``appropriate stimulus" around the critical period. Inspired by humans' early cognitive-developmental stage, we use multi-stage guidance near the critical period, and demonstrate the appropriate shaping reward (stage-2 guidance) in terms of the AI agent's performance, efficiency, and stability.

cs.AI

Toddler-Guidance Learning: Impacts of Critical Period on Multimodal AI Agents

Critical periods are phases during which a toddler's brain develops in spurts. To promote children's cognitive development, proper guidance is critical in this stage. However, it is not clear whether such a critical period also exists for the training of AI agents. Similar to human toddlers, well-timed guidance and multimodal interactions might significantly enhance the training efficiency of AI agents as well. To validate this hypothesis, we adapt this notion of critical periods to learning in AI agents and investigate the critical period in the virtual environment for AI agents. We formalize the critical period and Toddler-guidance learning in the reinforcement learning (RL) framework. Then, we built up a toddler-like environment with VECA toolkit to mimic human toddlers' learning characteristics. We study three discrete levels of mutual interaction: weak-mentor guidance (sparse reward), moderate mentor guidance (helper-reward), and mentor demonstration (behavioral cloning). We also introduce the EAVE dataset consisting of 30,000 real-world images to fully reflect the toddler's viewpoint. We evaluate the impact of critical periods on AI agents from two perspectives: how and when they are guided best in both uni- and multimodal learning. Our experimental results show that both uni- and multimodal agents with moderate mentor guidance and critical period on 1 million and 2 million training steps show a noticeable improvement. We validate these results with transfer learning on the EAVE dataset and find the performance advancement on the same critical period and the guidance.

cs.LG

VECA : A Toolkit for Building Virtual Environments to Train and Test Human-like Agents

Building human-like agent, which aims to learn and think like human intelligence, has long been an important research topic in AI. To train and test human-like agents, we need an environment that imposes the agent to rich multimodal perception and allows comprehensive interactions for the agent, while also easily extensible to develop custom tasks. However, existing approaches do not support comprehensive interaction with the environment or lack variety in modalities. Also, most of the approaches are difficult or even impossible to implement custom tasks. In this paper, we propose a novel VR-based toolkit, VECA, which enables building fruitful virtual environments to train and test human-like agents. In particular, VECA provides a humanoid agent and an environment manager, enabling the agent to receive rich human-like perception and perform comprehensive interactions. To motivate VECA, we also provide 24 interactive tasks, which represent (but are not limited to) four essential aspects in early human development: joint-level locomotion and control, understanding contexts of objects, multimodal learning, and multi-agent learning. To show the usefulness of VECA on training and testing human-like learning agents, we conduct experiments on VECA and show that users can build challenging tasks for engaging human-like algorithms, and the features supported by VECA are critical on training human-like agents.

cs.AI

Learning task-agnostic representation via toddler-inspired learning

One of the inherent limitations of current AI systems, stemming from the passive learning mechanisms (e.g., supervised learning), is that they perform well on labeled datasets but cannot deduce knowledge on their own. To tackle this problem, we derive inspiration from a highly intentional learning system via action: the toddler. Inspired by the toddler's learning procedure, we design an interactive agent that can learn and store task-agnostic visual representation while exploring and interacting with objects in the virtual environment. Experimental results show that such obtained representation was expandable to various vision tasks such as image classification, object localization, and distance estimation tasks. In specific, the proposed model achieved 100%, 75.1% accuracy and 1.62% relative error, respectively, which is noticeably better than autoencoder-based model (99.7%, 66.1%, 1.95%), and also comparable with those of supervised models (100%, 87.3%, 0.71%).

cs.AI

Algorithmic decomposition for efficient multiple nuclear spin detection in diamond

Efficiently detecting and characterizing individual spins in solid-state hosts is an essential step to expand the fields of quantum sensing and quantum information processing. While selective detection and control of a few 13C nuclear spins in diamond have been demonstrated using the electron spin of nitrogen-vacancy (NV) centers, a reliable, efficient, and automatic characterization method is desired. Here, we develop an automated algorithmic method for decomposing spectral data to identify and characterize multiple nuclear spins in diamond. We demonstrate efficient nuclear spin identification and accurate reproduction of hyperfine interaction components for both virtual and experimental nuclear spectroscopy data. We conduct a systematic analysis of this methodology and discuss the range of hyperfine interaction components of each nuclear spin that the method can efficiently detect. The result demonstrates a systematic approach that automatically detects nuclear spins with the aid of computational methods, facilitating the future scalability of devices.

quant-ph

Deep learning enhanced individual nuclear-spin detection

The detection of nuclear spins using individual electron spins has enabled new opportunities in quantum sensing and quantum information processing. Proof-of-principle experiments have demonstrated atomic-scale imaging of nuclear-spin samples and controlled multi-qubit registers. However, to image more complex samples and to realize larger-scale quantum processors, computerized methods that efficiently and automatically characterize spin systems are required. Here, we realize a deep learning model for automatic identification of nuclear spins using the electron spin of single nitrogen-vacancy (NV) centers in diamond as a sensor. Based on neural network algorithms, we develop noise recovery procedures and training sequences for highly non-linear spectra. We apply these methods to experimentally demonstrate fast identification of 31 nuclear spins around a single NV center and accurately determine the hyperfine parameters. Our methods can be extended to larger spin systems and are applicable to a wide range of electron-nuclear interaction strengths. These results enable efficient imaging of complex spin samples and automatic characterization of large spin-qubit registers.

quant-ph