SearcharxivSearch

arXiv subjects

Kyungmin Kim

Publications and source records attributed to Kyungmin Kim.

At least 19 recordsLinked to original sources

Iterative Projection-Based Embedding Scheme Combined with Variational Quantum Eigensolver

Quantum embedding methods offer a promising route to extend quantum chemical calculations to large multiscale systems by treating a chemically important subsystem at a high level of theory while describing its surrounding environment at an affordable level. The methods are also quite relevant for quantum computing approaches based on hardware with limited resources. Here, we present an iterative projection-based embedding framework combined with VQE, in which the environment density is allowed to respond self-consistently to the refined electronic structure of the embedded subsystem described by VQE. Unlike conventional one-shot approaches where the environment remains frozen after the initial orbital optimization, the proposed iterative scheme alternates between the VQE-level treatment of the subsystem and a mean-field-level refinement of the environment until mutual self-consistency is achieved. The convergence behavior of the scheme is first examined using several small test systems. Its practical applicability is then demonstrated with a composite system with a CH2NH molecule sandwiched by two benzene rings, with the C=N dihedral angle rotating from 0 to 90 deg. The iterative procedure consistently converges within ~10 iteration steps across all tested geometries, yielding energies below the conventional one-shot embedding results. The converged results well reproduce the fully correlated reference energy employing the same active space, and the resulting potential energy surface with respect to the dihedral rotation is also in good agreement with the reference one. These results demonstrate that our iterative embedding framework is numerically robust and physically sound, yielding a self-consistent and reliable treatment of inter-subsystem correlation. We expect that its formulation will be particularly compatible with the emerging paradigm of quantum-classical hybrid computing.

physics.chem-ph

Reconstruction-Dependent Imaging, Reactivity and Local Reduction of the CeO$_2$(100) surface

The possibility of mapping the local reactivity and reduction state to the atomic structure of chemically active oxide surfaces opens new avenues for further understanding of catalysis. Here, we combine scanning tunnelling (STM) and atomic force microscopy (AFM) with first-principles modelling to explore this possibility on the CeO2(100) surface. While STM reveals the periodicity of cerium-terminated and oxygen-terminated CeO$_2$(100) reconstructions coexisting on the same surface, AFM imaging and force spectroscopy provide direct identification of the exposed atomic species and their reactivity as the chemical interaction with the probe. Density functional theory based STM and AFM simulations reproduce the main experimental observations and show that STM contrast cannot be in general assigned to the atomic positions of certain chemical species, as traditionally assumed from previous studies. Simulated STM contrast of the two reconstructions across different reduction states associated with the removal of oxygen atoms in deeper layers, evidence that STM alone does not offer a robust fingerprint of the local reduction state for the cerium-terminated reconstruction, but it is sensitive to the reduced state in the case of the oxygen-terminated one, being able to provide information on a mixed distribution of Ce$^{3+}$ and Ce$^{4+}$ ions on the first sub-surface Ce layer.

cond-mat.mtrl-sci

The Interplay of Harness Design and Post-Training in LLM Agents

Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and what auxiliary information accompanies each per-step observation. While agents are routinely post-trained, this scaffolding is typically treated as a fixed engineering detail, with design effort limited to the training-free regime. Moreover, existing post-training algorithms assume a static environment, even though tool environments and tasks often shift upon deployment. To address this gap, we extend $\texttt{ALFWorld}$ (i) to treat the harness as a controllable design dimension and (ii) to support evaluation under task and tool environment shifts. Building on this, we systematically analyze how the harness design influences post-training in both in-distribution and out-of-distribution (OOD) settings. We empirically show that harness-aware post-training not only improves in-distribution performance but also enables agents to robustly adapt to OOD settings. Under a harness with minimal design effort, post-training suffers a drastic performance drop under stronger tool environment shifts, further highlighting the importance of harness-aware post-training under such shifts.

cs.LG

Verifiable Foundation Models for Robot Safety

Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes these models opaque and difficult to analyze formally, rendering them intractable for existing verification tools. In this paper, we present FEARL (Foundation-Enabled Assured Robot Learning), a framework that addresses this tension through a modular architectural decomposition. FEARL separates the policy into a large Controller (C) responsible for high-dimensional perception and task reasoning, and a small Safety module (S) that receives low-dimensional observations from dedicated safety sensors together with a bounded context embedding from C and produces the final action. Since many robot safety requirements, such as collision avoidance and workspace boundary constraints, can be expressed over these safety sensor observations, formal verification can be applied to S rather than to the full foundation-model backbone. This makes formal analysis tractable with existing tools while preserving the Controller's expressive power for task reasoning. To show that the decomposed policy remains capable of solving diverse tasks, we evaluate FEARL on three simulated robotic domains using multiple Controller backbones and training procedures, including pretrained off-the-shelf vision-language-action models. We further transfer the learned policy from one of our simulated tasks to a physical robot, suggesting that the low-dimensional safety interface supports practical sim-to-real transfer.

cs.RO

Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization

Uncertainty quantification is crucial in safety-critical systems, where decisions must be made under uncertainty. In particular, we consider the problem of online uncertainty quantification, where data points arrive sequentially. Online conformal prediction is a principled online uncertainty quantification method that dynamically constructs a prediction set at each time step. While existing methods for online conformal prediction provide long-run coverage guarantees without any distributional assumptions, they typically assume a full feedback setting in which the true label is always observed. In this paper, we propose a novel learning method for online conformal prediction with partial feedback from an adaptive adversary-a more challenging setup where the true label is revealed only when it lies inside the constructed prediction set. Specifically, we formulate online conformal prediction as an adversarial bandit problem by treating each candidate prediction set as an arm. Building on an existing algorithm for adversarial bandits, our method achieves a long-run coverage guarantee by explicitly establishing its connection to the regret of the learner. Finally, we empirically demonstrate the effectiveness of our method in both independent and identically distributed (i.i.d.) and non-i.i.d. settings, showing that it successfully controls the miscoverage rate while maintaining a reasonable size of the prediction set.

cs.LG

The impact of strong lensing on Hubble constant measurements with gravitational-wave dark sirens

The disagreement between early and late Universe electromagnetic measurements of the Hubble constant, $H_0$, known as the Hubble tension, highlights the need for independent and complementary probes. Gravitational-wave events have recently emerged as such a probe for constraining cosmological parameters. $H_{0}$ inference using these events relies on sky localization and luminosity distance estimates, both of which can be significantly improved for strongly lensed events with appropriate lens modeling. In this context, we propose utilizing strong lensing of dark sirens, gravitational-wave events without identified electromagnetic counterparts, in combination with strong lensing of galaxies as a novel method for measuring $H_0$. The constant is inferred from the luminosity distances of these lensed dark sirens and the redshifts of their host galaxies, combining information from individual events to obtain statistically stronger constraints when multiple events are available. We adopt a simulated galaxy catalog, \texttt{MICECATv2}, as the basis for simulating strong lensing of galaxies and to provide the redshift information of host galaxy candidates required to infer $H_0$. We also examine the impact of galaxy catalog incompleteness on the resulting $H_0$ inference. Our results demonstrate that using only 8 strongly lensed dark sirens, analyzed with a dedicated galaxy-galaxy lensing catalog, can improve the precision of $H_{0}$ by roughly 50\% compared to 250 unlensed events.

astro-ph.CO

Transductive Generalization via Optimal Transport and Its Application to Graph Node Classification

Many existing transductive bounds rely on classical complexity measures that are computationally intractable and often misaligned with empirical behavior. In this work, we establish new representation-based generalization bounds in a distribution-free transductive setting, where learned representations are dependent, and test features are accessible during training. We derive global and class-wise bounds via optimal transport, expressed in terms of Wasserstein distances between encoded feature distributions. We demonstrate that our bounds are efficiently computable and strongly correlate with empirical generalization in graph node classification, improving upon classical complexity measures. Additionally, our analysis reveals how the GNN aggregation process transforms the representation distributions, inducing a trade-off between intra-class concentration and inter-class separation. This yields depth-dependent characterizations that capture the non-monotonic relationship between depth and generalization error observed in practice. The code is available at https://github.com/ml-postech/Transductive-OT-Gen-Bound.

cs.LG

Vanishing Compactness Gap and Fermionic Compact Dark Matter in Ho\v{r}ava-Lifshitz Gravity

We show that the gap in the compactness between black holes and neutron stars witnessed in general relativity may be vanishing in Ho\v{r}ava-Lifshitz (HL) gravity. Assuming a fermion equation-of-state for simplicity, and solving the Tolman-Oppenheimer-Volkoff equation within the HL gravity framework, we see that there exists a minimum fermion mass $m_f^\text{(min)}(q,y)$, above which the gap of the compactness between black hole and fermionic compact object vanishes, for a given deformation parameter $q$ of HL and interaction strength $y$ between fermions. Thus, in HL gravity, the mass and radius of an object found in the lower mass gap by LIGO-Virgo-KAGRA observations might not be able to classify it as a black hole or a neutron star. It is interesting to note that a fermion of mass $\sim 40\ \text{GeV}$ can form a highly compact object of mass $\sim 10^{-4}\ \msun$ and radius $\sim 1\ \text{m}$ that may play the role of the cold dark matter. In addition, we find the possible existence of another class of compact objects whose compactness is comparable to that of a black hole.

gr-qc

Extending the Handover-Iterative VQE to Challenging Strongly Correlated Systems: $N_2$ and Fe-S Cluster

Accurately describing strongly correlated electronic systems remains a central challenge in quantum chemistry, as electron-electron interactions give rise to complex many-body wavefunctions that are difficult to capture with conventional approximations. Classical wavefunction-based approaches, such as the Semistochastic Heat-bath Configuration Interaction (SHCI) and the Density Matrix Renormalization Group (DMRG), currently define the state of the art, systematically converging toward the Full Configuration Interaction (FCI) limit, but at a rapidly increasing computational cost. Quantum computing algorithms promise to alleviate this scaling bottleneck by leveraging entanglement and superposition to represent correlated states more compactly. We introduced the Handover-Iterative Variational Quantum Eigensolver (HI-VQE) as a practical quantum computing algorithm with an iterative "handover" mechanism that dynamically exchanges information between quantum and classical computers, even using Noisy Intermediate-Scale Quantum (NISQ) computers. In this work, we extend the HI-VQE to benchmark two prototypical strongly correlated systems, the nitrogen molecule $N_2$ and iron-sulfur (Fe-S) cluster, which serve as stringent tests for both classical and quantum electronic-structure methods. By comparing HI-VQE results against Heat-bath Configuration Interaction (HCI) benchmarks, we assess its accuracy, scalability, and ability to capture multireference correlation effects. Achieving quantitative agreement on these canonical systems demonstrates a viable pathway toward quantum-enhanced simulations of complex bioinorganic molecules, catalytic mechanisms, and correlated materials.

quant-ph

Probabilistic Multi-Agent Aircraft Landing Time Prediction

Accurate and reliable aircraft landing time prediction is essential for effective resource allocation in air traffic management. However, the inherent uncertainty of aircraft trajectories and traffic flows poses significant challenges to both prediction accuracy and trustworthiness. Therefore, prediction models should not only provide point estimates of aircraft landing times but also the uncertainties associated with these predictions. Furthermore, aircraft trajectories are frequently influenced by the presence of nearby aircraft through air traffic control interventions such as radar vectoring. Consequently, landing time prediction models must account for multi-agent interactions in the airspace. In this work, we propose a probabilistic multi-agent aircraft landing time prediction framework that provides the landing times of multiple aircraft as distributions. We evaluate the proposed framework using an air traffic surveillance dataset collected from the terminal airspace of the Incheon International Airport in South Korea. The results demonstrate that the proposed model achieves higher prediction accuracy than the baselines and quantifies the associated uncertainties of its outcomes. In addition, the model uncovered underlying patterns in air traffic control through its attention scores, thereby enhancing explainability.

cs.MA

Model-Based Reinforcement Learning under Random Observation Delays

Delays frequently occur in real-world environments, yet standard reinforcement learning (RL) algorithms often assume instantaneous perception of the environment. We study random sensor delays in POMDPs, where observations may arrive out-of-sequence, a setting that has not been previously addressed in RL. We analyze the structure of such delays and demonstrate that naive approaches, such as stacking past observations, are insufficient for reliable performance. To address this, we propose a model-based filtering process that sequentially updates the belief state based on an incoming stream of observations. We then introduce a simple delay-aware framework that incorporates this idea into model-based RL, enabling agents to effectively handle random delays. Applying this framework to the Dreamer world-modeling scheme, our method consistently outperforms delay-aware baselines developed for MDPs and demonstrates robustness to delay distribution shifts during deployment. Additionally, we present experiments on simulated robotic tasks, comparing our method to common practical heuristics and emphasizing the importance of explicitly modeling observation delays.

cs.LG

Near-surface Defects Break Symmetry in Water Adsorption on CeO$_{2-x}$(111)

Water interactions with oxygen-deficient cerium dioxide (CeO$_2$) surfaces are central to hydrogen production and catalytic redox reactions, but the atomic-scale details of how defects influence adsorption and reactivity remain elusive. Here, we unveil how water adsorbs on partially reduced CeO$_{2-x}$(111) using atomic force microscopy (AFM) with chemically sensitive, oxygen-terminated probes, combined with first-principles calculations. Our AFM imaging reveals water molecules as sharp, asymmetric boomerang-like features radically departing from the symmetric triangular motifs previously attributed to molecular water. Strikingly, these features localize near subsurface defects. While the experiments are carried out at cryogenic temperature, water was dosed at room temperature, capturing configurations relevant to initial adsorption events in catalytic processes. Density functional theory identifies Ce$^{3+}$ sites adjacent to subsurface vacancies as the thermodynamically favored adsorption sites, where defect-induced symmetry breaking governs water orientation. Force spectroscopy and simulations further distinguish Ce$^{3+}$ from Ce$^{4+}$ centers through their unique interaction signatures. By resolving how subsurface defects control water adsorption at the atomic scale, this work demonstrates the power of chemically selective AFM for probing site-specific reactivity in oxide catalysts, laying the groundwork for direct investigations of complex systems such as single-atom catalysts, metal-support interfaces, and defect-engineered oxides.

cond-mat.mtrl-sci

Adapting World Models with Latent-State Dynamics Residuals

Simulation-to-reality reinforcement learning (RL) faces the critical challenge of reconciling discrepancies between simulated and real-world dynamics, which can severely degrade agent performance. A promising approach involves learning corrections to simulator forward dynamics represented as a residual error function, however this operation is impractical with high-dimensional states such as images. To overcome this, we propose ReDRAW, a latent-state autoregressive world model pretrained in simulation and calibrated to target environments through residual corrections of latent-state dynamics rather than of explicit observed states. Using this adapted world model, ReDRAW enables RL agents to be optimized with imagined rollouts under corrected dynamics and then deployed in the real world. In multiple vision-based MuJoCo domains and a physical robot visual lane-following task, ReDRAW effectively models changes to dynamics and avoids overfitting in low data regimes where traditional transfer methods fail.

cs.LG

Selective Generation for Controllable Language Models

Trustworthiness of generative language models (GLMs) is crucial in their deployment to critical decision making systems. Hence, certified risk control methods such as selective prediction and conformal prediction have been applied to mitigating the hallucination problem in various supervised downstream tasks. However, the lack of appropriate correctness metric hinders applying such principled methods to language generation tasks. In this paper, we circumvent this problem by leveraging the concept of textual entailment to evaluate the correctness of the generated sequence, and propose two selective generation algorithms which control the false discovery rate with respect to the textual entailment relation (FDR-E) with a theoretical guarantee: $\texttt{SGen}^{\texttt{Sup}}$ and $\texttt{SGen}^{\texttt{Semi}}$. $\texttt{SGen}^{\texttt{Sup}}$, a direct modification of the selective prediction, is a supervised learning algorithm which exploits entailment-labeled data, annotated by humans. Since human annotation is costly, we further propose a semi-supervised version, $\texttt{SGen}^{\texttt{Semi}}$, which fully utilizes the unlabeled data by pseudo-labeling, leveraging an entailment set function learned via conformal prediction. Furthermore, $\texttt{SGen}^{\texttt{Semi}}$ enables to use more general class of selection functions, neuro-selection functions, and provides users with an optimal selection function class given multiple candidates. Finally, we demonstrate the efficacy of the $\texttt{SGen}$ family in achieving a desired FDR-E level with comparable selection efficiency to those from baselines on both open and closed source GLMs. Code and datasets are provided at https://github.com/ml-postech/selective-generation.

cs.LG

TARDiS : Text Augmentation for Refining Diversity and Separability

Text augmentation (TA) is a critical technique for text classification, especially in few-shot settings. This paper introduces a novel LLM-based TA method, TARDiS, to address challenges inherent in the generation and alignment stages of two-stage TA methods. For the generation stage, we propose two generation processes, SEG and CEG, incorporating multiple class-specific prompts to enhance diversity and separability. For the alignment stage, we introduce a class adaptation (CA) method to ensure that generated examples align with their target classes through verification and modification. Experimental results demonstrate TARDiS's effectiveness, outperforming state-of-the-art LLM-based TA methods in various few-shot text classification tasks. An in-depth analysis confirms the detailed behaviors at each stage.

cs.CL

Realizable Continuous-Space Shields for Safe Reinforcement Learning

While Deep Reinforcement Learning (DRL) has achieved remarkable success across various domains, it remains vulnerable to occasional catastrophic failures without additional safeguards. An effective solution to prevent these failures is to use a shield that validates and adjusts the agent's actions to ensure compliance with a provided set of safety specifications. For real-world robotic domains, it is essential to define safety specifications over continuous state and action spaces to accurately account for system dynamics and compute new actions that minimally deviate from the agent's original decision. In this paper, we present the first shielding approach specifically designed to ensure the satisfaction of safety requirements in continuous state and action spaces, making it suitable for practical robotic applications. Our method builds upon realizability, an essential property that confirms the shield will always be able to generate a safe action for any state in the environment. We formally prove that realizability can be verified for stateful shields, enabling the incorporation of non-Markovian safety requirements, such as loop avoidance. Finally, we demonstrate the effectiveness of our approach in ensuring safety without compromising the policy's success rate by applying it to a navigation problem and a multi-agent particle environment.

cs.LG

Can we discern millilensed gravitational-wave signals from signals produced by precessing binary black holes with ground-based detectors?

Millilensed gravitational waves (GWs) can potentially be identified by the interference signatures caused by $\sim\!O(10\textrm{--}100)~\textrm{ms}$ time delays between multiple overlapping lensed signals. However, distinguishing millilensed GWs from GWs generated by precessing binary black-hole mergers can be challenging due to their apparent similar waveform shapes. This morphological similarity may be an obstacle to template-based searches to correctly identifying the origin of observed GWs and poses a fundamental question, can we discern millilensed GW signals from signals produced by precessing binary black holes? In this study, we investigate the feasibility of distinguishing between these GWs by performing a proof-of-principle injection study of simulated millilensed precessing GW signals, within the context of ground-based LIGO-Virgo-KAGRA detector network detections. Our findings indicate that it is possible to differentiate between the two effects by comparing signal-to-noise ratios (SNRs) computed using templates based on different hypotheses for the target signal. We further show from the parameter estimation study that while lensing magnification is sensitive to precession, it is possible to identify millilensing in precessing GW signals with an SNR of 18. The recovery of precession in the presence of lensing is more challenging but improves significantly for signals with an SNR of 40. Nonetheless, neglecting millilensing effects results in biases in the recovered spins, revealing the importance of accounting for these effects in accurate GW signal analysis.

gr-qc

Statistical Discrimination in Ratings-Guided Markets

We study statistical discrimination of individuals based on payoff-irrelevant social identities in markets that utilize ratings and recommendations for social learning. Even though rating/recommendation algorithms can be designed to be fair and unbiased, ratings-based social learning can still lead to discriminatory outcomes. Our model demonstrates how users' attention choices can result in asymmetric data sampling across social groups, leading to discriminatory inferences and potential discrimination based on group identities.

cs.GT