SearcharxivSearch

arXiv subjects

Ji-Hoon Park

Publications and source records attributed to Ji-Hoon Park.

16 recordsLinked to original sources

DQE-CIR: Distinctive Query Embeddings through Learnable Attribute Weights and Target Relative Negative Sampling in Composed Image Retrieval

Composed image retrieval (CIR) addresses the task of retrieving a target image by jointly interpreting a reference image and a modification text that specifies the intended change. Most existing methods are still built upon contrastive learning frameworks that treat the ground truth image as the only positive instance and all remaining images as negatives. This strategy inevitably introduces relevance suppression, where semantically related yet valid images are incorrectly pushed away, and semantic confusion, where different modification intents collapse into overlapping regions of the embedding space. As a result, the learned query representations often lack discriminativeness, particularly at fine-grained attribute modifications. To overcome these limitations, we propose distinctive query embeddings through learnable attribute weights and target relative negative sampling (DQE-CIR), a method designed to learn distinctive query embeddings by explicitly modeling target relative relevance during training. DQE-CIR incorporates learnable attribute weighting to emphasize distinctive visual features conditioned on the modification text, enabling more precise feature alignment between language and vision. Furthermore, we introduce target relative negative sampling, which constructs a target relative similarity distribution and selects informative negatives from a mid-zone region that excludes both easy negatives and ambiguous false negatives. This strategy enables more reliable retrieval for fine-grained attribute changes by improving query discriminativeness and reducing confusion caused by semantically similar but irrelevant candidates.

cs.CV

Three-dimensional imaging of individual carbon atoms

Carbon is fundamental to science and technology due to its diverse bonding configurations, structural versatility, and essential role in defining the mechanical, chemical, electronic, and quantum properties of materials. However, direct three-dimensional (3D) imaging of individual carbon atoms remains a long-standing challenge. Here, we use twisted bilayer graphene (TBG) as a model system and demonstrate ptychographic atomic electron tomography (pAET) for determining the 3D atomic coordinates of individual carbon atoms with a precision of 0.11 angstrom. The resulting 3D atomic model uncovers chiral lattice distortions driven by van der Waals interactions that exhibit meron-like and skyrmion-like structural textures. These findings provide direct insight into the interplay between 3D chiral lattice deformation and electronic properties in moire-engineered carbon systems. Beyond TBG, pAET offers a versatile approach for 3D atomic-scale imaging of carbon-based and other light-element materials that are central to advances in physics, chemistry, materials science, and nanotechnology.

cond-mat.mtrl-sci

Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models

Current video understanding models excel at recognizing "what" is happening but fall short in high-level cognitive tasks like causal reasoning and future prediction, a limitation rooted in their lack of commonsense world knowledge. To bridge this cognitive gap, we propose a novel framework that synergistically fuses a powerful Vision Foundation Model (VFM) for deep visual perception with a Large Language Model (LLM) serving as a knowledge-driven reasoning core. Our key technical innovation is a sophisticated fusion module, inspired by the Q-Former architecture, which distills complex spatiotemporal and object-centric visual features into a concise, language-aligned representation. This enables the LLM to effectively ground its inferential processes in direct visual evidence. The model is trained via a two-stage strategy, beginning with large-scale alignment pre-training on video-text data, followed by targeted instruction fine-tuning on a curated dataset designed to elicit advanced reasoning and prediction skills. Extensive experiments demonstrate that our model achieves state-of-the-art performance on multiple challenging benchmarks. Notably, it exhibits remarkable zero-shot generalization to unseen reasoning tasks, and our in-depth ablation studies validate the critical contribution of each architectural component. This work pushes the boundary of machine perception from simple recognition towards genuine cognitive understanding, paving the way for more intelligent and capable AI systems in robotics, human-computer interaction, and beyond.

cs.CV

Multi-stage Prompt Refinement for Mitigating Hallucinations in Large Language Models

Recent advancements in large language models (LLMs) have shown strong performance in natural language understanding and generation tasks. However, LLMs continue to encounter challenges with hallucinations, where models generate plausible but incorrect information. While several factors contribute to hallucinations, the impact of ill-formed prompts, prompts with ambiguous wording, incorrect grammar, or incomplete information, was relatively under explored. To address this, we introduce Multi-stage Prompt Refinement (MPR), a framework designed to systematically improve these ill-formed prompts across multiple stages. Each stage addresses specific errors such as punctuation, typographical mistakes, and misuse of key terms, using small language models (SLMs) fine-tuned for these tasks. MPR iteratively enhances the clarity of prompts with additional context and employs a self-reflection mechanism with ranking to prioritize the most relevant input. Experimental results on hallucination benchmarks show that prompts refined by MPR achieve over an 85~\% win rate compared to their original forms, demonstrating its effectiveness in reducing hallucinations and improving LLM output accuracy. Interestingly, we reveal that MPR can be combined with existing post-hoc hallucination mitigation frameworks, further enhancing its versatility. MPR provides a lightweight and adaptable solution for enhancing LLM reliability across various domains.

cs.CL

CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement

Recent advancements in large language models (LLMs) highlight their fluency in generating responses to diverse prompts. However, these models sometimes generate plausible yet incorrect ``hallucinated" facts, undermining trust. A frequent but often overlooked cause of such errors is the use of poorly structured or vague prompts by users, leading LLMs to base responses on assumed rather than actual intentions. To mitigate hallucinations induced by these ill-formed prompts, we introduce Curative Prompt Refinement (CPR), a plug-and-play framework for curative prompt refinement that 1) cleans ill-formed prompts, and 2) generates additional informative task descriptions to align the intention of the user and the prompt using a fine-tuned small language model. When applied to language models, we discover that CPR significantly increases the quality of generation while also mitigating hallucination. Empirical studies show that prompts with CPR applied achieves over a 90\% win rate over the original prompts without any external knowledge.

cs.CL

KiC: Keyword-inspired Cascade for Cost-Efficient Text Generation with LLMs

Large language models (LLMs) have demonstrated state-of-the-art performance across a wide range of natural language processing tasks. However, high-performing models are typically accessible only via APIs, incurring substantial inference costs. Cascade methods address this by initially employing a cheaper model and escalating to a stronger one only when necessary. Nevertheless, existing cascade approaches struggle to select a reliable representative response and assess the overall reliability of free-form outputs, as they rely on exact text matching. To overcome these limitations, we propose Keyword-inspired Cascade (KiC), a novel framework for cost-efficient free-form text generation. KiC identifies the most representative answer among multiple outputs from a weaker model and evaluates the semantic alignment of other responses with it. Based on the degree of alignment, KiC determines whether to accept the weaker model's output or escalate to a stronger model. Experiments on three free-form text generation benchmarks show that KiC achieves 97.53 percent of GPT-4's accuracy while reducing API costs by 28.81 percent on average, and even outperforms GPT-4 in a specific benchmark.

cs.CL

Bolometric Superconducting Optical Nanoscopy (BOSON)

Superconducting transition-edge sensors are renowned for their extraordinary photon sensitivity and energy resolution, finding applications spanning quantum information, astronomy, and nanophotonics. Here, we report the development of BOlometric Superconducting Optical Nanoscopy (BOSON), a novel platform that integrates bolometric detection at the superconducting transition edges with near-field optical techniques. BOSON enables the mapping of photoinduced changes in superconductivity with unprecedented spatial resolution and photon sensitivity. By incorporating BOSON with low-dimensional materials, we achieved polariton imaging at nanowatt excitation levels--at least four orders of magnitude lower than the power typically required in prior near-field nanoscopy experiments. Our findings highlight the potential for BOSON to advance scanning probe based optical platforms to enable the detection of photons, polaritons, and Cooper pair dynamics at the nanoscale. This paves the way for quantum sensing applications using single-polariton detection and can offer deeper insights into quasiparticle dynamics.

quant-ph

GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs

Large Language Models (LLMs) trained on extensive datasets often learn sensitive information, which raises significant social and legal concerns under principles such as the "Right to be forgotten." Retraining entire models from scratch to remove undesired information is both costly and impractical. Furthermore, existing single-domain unlearning methods fail to address multi-domain scenarios, where knowledge is interwoven across domains such as privacy and copyright, creating overlapping representations that lead to excessive knowledge removal or degraded performance. To tackle these issues, we propose GRAIL (GRadient-based AdaptIve unLearning), a novel multi-domain unlearning framework. GRAIL leverages gradient information from multiple domains to precisely distinguish the unlearning scope from the retention scope, and applies an adaptive parameter-wise localization strategy to selectively remove targeted knowledge while preserving critical parameters for each domain. Experimental results on unlearning benchmarks show that GRAIL achieves unlearning success on par with the existing approaches, while also demonstrating up to 17% stronger knowledge retention success compared to the previous state-of-art method. Our findings establish a new paradigm for effectively managing and regulating sensitive information in large-scale pre-trained language models.

cs.CL

SUGAR: Leveraging Contextual Confidence for Smarter Retrieval

Bearing in mind the limited parametric knowledge of Large Language Models (LLMs), retrieval-augmented generation (RAG) which supplies them with the relevant external knowledge has served as an approach to mitigate the issue of hallucinations to a certain extent. However, uniformly retrieving supporting context makes response generation source-inefficient, as triggering the retriever is not always necessary, or even inaccurate, when a model gets distracted by noisy retrieved content and produces an unhelpful answer. Motivated by these issues, we introduce Semantic Uncertainty Guided Adaptive Retrieval (SUGAR), where we leverage context-based entropy to actively decide whether to retrieve and to further determine between single-step and multi-step retrieval. Our empirical results show that selective retrieval guided by semantic uncertainty estimation improves the performance across diverse question answering tasks, as well as achieves a more efficient inference.

cs.CL

Explaining generative diffusion models via visual analysis for interpretable decision-making process

Diffusion models have demonstrated remarkable performance in generation tasks. Nevertheless, explaining the diffusion process remains challenging due to it being a sequence of denoising noisy images that are difficult for experts to interpret. To address this issue, we propose the three research questions to interpret the diffusion process from the perspective of the visual concepts generated by the model and the region where the model attends in each time step. We devise tools for visualizing the diffusion process and answering the aforementioned research questions to render the diffusion process human-understandable. We show how the output is progressively generated in the diffusion process by explaining the level of denoising and highlighting relationships to foundational visual concepts at each time step through the results of experiments with various visual analyses using the tools. Throughout the training of the diffusion model, the model learns diverse visual concepts corresponding to each time-step, enabling the model to predict varying levels of visual concepts at different stages. We substantiate our tools using Area Under Cover (AUC) score, correlation quantification, and cross-attention mapping. Our findings provide insights into the diffusion process and pave the way for further research into explainable diffusion mechanisms.

cs.CV

NeuroInspect: Interpretable Neuron-based Debugging Framework through Class-conditional Visualizations

Despite deep learning (DL) has achieved remarkable progress in various domains, the DL models are still prone to making mistakes. This issue necessitates effective debugging tools for DL practitioners to interpret the decision-making process within the networks. However, existing debugging methods often demand extra data or adjustments to the decision process, limiting their applicability. To tackle this problem, we present NeuroInspect, an interpretable neuron-based debugging framework with three key stages: counterfactual explanations, feature visualizations, and false correlation mitigation. Our debugging framework first pinpoints neurons responsible for mistakes in the network and then visualizes features embedded in the neurons to be human-interpretable. To provide these explanations, we introduce CLIP-Illusion, a novel feature visualization method that generates images representing features conditioned on classes to examine the connection between neurons and the decision layer. We alleviate convoluted explanations of the conventional visualization approach by employing class information, thereby isolating mixed properties. This process offers more human-interpretable explanations for model errors without altering the trained network or requiring additional data. Furthermore, our framework mitigates false correlations learned from a dataset under a stochastic perspective, modifying decisions for the neurons considered as the main causes. We validate the effectiveness of our framework by addressing false correlations and improving inferences for classes with the worst performance in real-world settings. Moreover, we demonstrate that NeuroInspect helps debug the mistakes of DL models through evaluation for human understanding. The code is openly available at https://github.com/yeongjoonJu/NeuroInspect.

cs.CV

Tuning color centers at a twisted interface

Color center is a promising platform for quantum technologies, but their application is hindered by the typically random defect distribution and complex mesoscopic environment. Employing cathodoluminescence, we demonstrate that an ultraviolet-emitting single photon emitter can be readily activated and controlled on-demand at the twisted interface of two hexagonal boron nitride flakes. The brightness of the color center can be enhanced by two orders of magnitude by altering the twist angle. Additionally, a brightness modulation of nearly 100% of this color center is achieved by an external voltage. Our ab-initio GW calculations suggest that the emission is correlated to nitrogen vacancies and that a twist-induced moiré potential facilitates electron-hole recombination. This mechanism is further exploited to draw nanoscale color center patterns using electron beams.

cond-mat.mtrl-sci

Epitaxial Growth of a Single-Crystal Hybridized Boron Nitride and Graphene layer on a Wide-Band Gap Semiconductor

Vertical and lateral heterogeneous structures of two-dimensional (2D) materials have paved the way for pioneering studies on the physics and applications of 2D materials. A hybridized hexagonal boron nitride (h-BN) and graphene lateral structure, a heterogeneous 2D structure, has been fabricated on single-crystal metals or metal foils by chemical vapor deposition (CVD). However, once fabricated on metals, the h-BN/graphene lateral structures require an additional transfer process for device applications, as reported for CVD graphene grown on metal foils. Here, we demonstrate that a single-crystal h-BN/graphene lateral structure can be epitaxially grown on a wide-gap semiconductor, SiC(0001). First, a single-crystal h-BN layer with the same orientation as bulk SiC was grown on a Si-terminated SiC substrate at 850 oC using borazine molecules. Second, when heated above 1150 oC in vacuum, the h-BN layer was partially removed and, subsequently, replaced with graphene domains. Interestingly, these graphene domains possess the same orientation as the h-BN layer, resulting in a single-crystal h-BN/graphene lateral structure on a whole sample area. For temperatures above 1600 oC, the single-crystal h-BN layer was completely replaced by the single-crystal graphene layer. The crystalline structure, electronic band structure, and atomic structure of the h-BN/graphene lateral structure were studied by using low energy electron diffraction, angle-resolved photoemission spectroscopy, and scanning tunneling microscopy, respectively. The h-BN/graphene lateral structure fabricated on a wide-gap semiconductor substrate can be directly applied to devices without a further transfer process, as reported for epitaxial graphene on a SiC substrate.

cond-mat.mes-hall

Designed Three-Dimensional Freestanding Single-Crystal Carbon Architectures

Single-crystal carbon nanomaterials have led to great advances in nanotechnology. The first single-crystal carbon nanomaterial, fullerene, was fabricated in a zero-dimensional form. One-dimensional carbon nanotubes and two-dimensional graphene have since followed and continue to provide further impetus to this field. In this study, we fabricated designed three-dimensional (3D) single-crystal carbon architectures by using silicon carbide templates. For this method, a designed 3D SiC structure was transformed into a 3D freestanding single-crystal carbon structure that retained the original SiC structure by performing a simple single-step thermal process. The SiC structure inside the 3D carbon structure is self-etched, which results in a 3D freestanding carbon structure. The 3D carbon structure is a single crystal with the same hexagonal close-packed structure as graphene. The size of the carbon structures can be controlled from the nanoscale to the microscale, and arrays of these structures can be scaled up to the wafer scale. The 3D freestanding carbon structures were found to be mechanically stable even after repeated loading. The relationship between the reversible mechanical deformation of a carbon structure and its electrical conductance was also investigated. Our method of fabricating designed 3D freestanding single-crystal graphene architectures opens up prospects in the field of single-crystal carbon nanomaterials, and paves the way for the development of 3D single-crystal carbon devices.

cond-mat.mes-hall

Opening and reversible control of a wide energy gap in uniform monolayer graphene

For graphene to be used in semiconductor applications, a wide energy gap of at least 0.5 eV at the Dirac energy must be opened without the introduction of atomic defects. However, such a wide energy gap has not been realized in graphene, except in the cases of narrow, chemically terminated graphene nanostructures with inevitable edge defects. Here, we demonstrated that a wide energy gap of 0.74 eV, which is larger than that of germanium, could be opened in uniform monolayer graphene without the introduction of atomic defects into graphene. The wide energy gap was opened through the adsorption of self-assembled twisted sodium nanostrips. Furthermore, the energy gap was reversibly controllable through the alternate adsorption of sodium and oxygen. The opening of such a wide energy gap with minimal degradation of mobility could improve the applicability of graphene in semiconductor devices, which would result in a major advancement in graphene technology.

cond-mat.mes-hall

Theory of magnetic enhancement in strontium hexaferrite through Zn-Sn pair substitution

We study the site occupancy and magnetic properties of Zn-Sn substituted M-type Sr-hexaferrite SrFe$_{12-x}$(Zn$_{0.5}$Sn$_{0.5}$)$_x$O$_{19}$ with x = 1 using first-principles total-energy calculations. We find that in a ground-state configuration Zn-Sn ions preferentially occupy $4f_1$ and $4f_2$ sites unlike the model previously suggested by Ghasemi et al. [J. Appl. Phys, \textbf{107}, 09A734 (2010)], where Zn$^{2+}$ and Sn$^{4+}$ ions occupy the $2b$ and $4f_2$ sites. Density-functional theory calculations show that our model has a lower total energy by more than 0.2 eV per unit cell compared to Ghasemi's model. More importantly, the latter does not show an increase in saturation magnetization ($M_s$) compared to the pure $M$-type Sr-hexaferrite, in disagreement with the experiment. On the other hand, our model correctly predicts a rapid increase in $M_s$ as well as a decrease in magnetic anisotropy compared to the pure $M$-type Sr-hexaferrite, consistent with experimental measurements.

cond-mat.mtrl-sci