SearcharxivSearch

arXiv subjects

Jeongmo Kim

Publications and source records attributed to Jeongmo Kim.

4 recordsLinked to original sources

Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning

Cross-domain offline reinforcement learning aims to adapt a policy from a source domain to a target domain using only pre-collected datasets, where environment dynamics may differ. A key challenge is to leverage source data while reducing distributional mismatch, particularly when the target dataset is extremely limited. To address this, we propose Target-aligned Coverage Expansion (TCE), a framework that decides how source data should be used, either by directly incorporating target-near transitions or by expanding state coverage through target-aligned generation, guided by theoretical analysis. TCE builds on a dual score-based generative model to synthesize target-consistent transitions over an expanded state region. Extensive experiments across diverse cross-domain environments show that TCE consistently outperforms state-of-the-art cross-domain offline RL baselines.

cs.LG

Strict Subgoal Execution: Reliable Long-Horizon Planning in Hierarchical Reinforcement Learning

Long-horizon goal-conditioned tasks pose fundamental challenges for reinforcement learning (RL), particularly when goals are distant and rewards are sparse. While hierarchical and graph-based methods offer partial solutions, their reliance on conventional hindsight relabeling often fails to correct subgoal infeasibility, leading to inefficient high-level planning. To address this, we propose Strict Subgoal Execution (SSE), a graph-based hierarchical RL framework that integrates Frontier Experience Replay (FER) to separate unreachable from admissible subgoals and streamline high-level decision making. FER delineates the reachability frontier using failure and partial-success transitions, which identifies unreliable subgoals, increases subgoal reliability, and reduces unnecessary high-level decisions. Additionally, SSE employs a decoupled exploration policy to cover underexplored regions of the goal space and a path refinement that adjusts edge costs using observed low-level failures. Experimental results across diverse long-horizon benchmarks show that SSE consistently outperforms existing goal-conditioned and hierarchical RL methods in both efficiency and success rate. Our code is available at https://jaebak1996.github.io/SSE/

cs.LG

Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks

Meta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git.

cs.LG

Charge-Driven Liquid-Crystalline Behavior of Ligand-Functionalized Nanorods in Apolar Solvent

Concentrated colloidal suspensions of nanorods often exhibit liquid-crystalline (LC) behavior. The transition to a nematic LC phase, with long-range orientational order of the particles, is usually well captured by Onsager's theory for hard rods, at least qualitatively. The theory shows how the volume fraction at the transition decreases with increasing aspect ratio of the rods. It also explains that the long-range electrostatic repulsive interaction occurring between rods stabilized by their surface charge can significantly increase their effective diameter, resulting in a decrease of the volume fraction at the transition, as compared to sterically stabilized rods. Here, we report on a system of ligand-stabilized LaPO4 nanorods, of aspect ratio around 11, dispersed in apolar medium exhibiting the counter-intuitive observation that the onset of nematic self-assembly occurs at an extremely low volume fraction of around 0.25%, which is lower than observed (around 3%) with the same particles when charge-stabilized in polar solvent. Furthermore, the nanorod volume fraction at the transition increases with increasing concentration of ligands, in a similar way as in polar media where increasing the ionic strength leads to surface-charge screening. This peculiar system was investigated by dynamic light scattering, Fourier-Transform Infra-Red spectroscopy, zetametry, electron microscopy, polarized-light microscopy, photoluminescence measurements, and X-ray scattering. Based on these experimental data, we formulate several tentative scenarios that might explain this unexpected phase behavior. However, at this stage, its full understanding remains a pending theoretical challenge. Nevertheless, this study shows that dispersing anisotropic nanoparticles in an apolar solvent may sometimes lead to spontaneous ordering events that defy our intuitive ideas about colloidal systems.

cond-mat.soft