SearcharxivSearch

arXiv subjects

Minjae Kwon

Publications and source records attributed to Minjae Kwon.

8 recordsLinked to original sources

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under out-of-distribution (OOD) deployment shifts, pretraining-only safe ICRL can give poor reward-safety tradeoffs because the remaining budget affects behavior only through frozen policy conditioning, not an explicit action-level check against predicted future cost. We propose a latent Q-Barrier shield that learns a context representation, latent dynamics, and an ensemble cost critic before deployment. Without parameter updates, the shield infers context from history and filters or softly reweights candidate actions using the remaining budget and predicted future cost. We prove a conditional, error-decomposed barrier-margin result: a Q-Barrier-satisfying action leaves the next latent-budget state with an approximately budget-safe continuation under the learned critic, up to Bellman and latent-prediction errors. Across five safe ICRL benchmarks, the shield improves deployment-time reward-safety tradeoffs over a strong safe-ICRL baseline: after a short context window, it achieves higher return in four of five benchmarks while matching or lowering average episode cost in all five.

cs.LG

Cold Neutron Imaging and Efficiency Measurements with a Boron-10 Coated Double-GEM Detector

A ${}^{10}\mathrm{B}$-coated double-GEM neutron detector (BGEM) was developed as a ${}^{3}\mathrm{He}$-free cold-neutron beamline detector using a single $\mathrm{B}_{4}\mathrm{C}$ converter cathode and a 512-channel APV25 orthogonal-strip readout over an active area of $10 \times 10~\mathrm{cm}^{2}$. The detector was tested at the HANARO Bio-REF beamline with a monochromatic $4.5~\mathring{\mathrm{A}}$ beam ($E_{n}=4.03~\mathrm{meV}$). The absolute detection efficiency relative to a ${}^{6}\mathrm{Li}$-based Ce:LiCAF reference detector was $\varepsilon_{\mathrm{BGEM}}=(8.69 \pm 0.20)\%$ (stat.). The pulse-height spectrum was qualitatively consistent with Geant4 energy-deposition simulations, and Cd-mask imaging yielded a Gaussian-equivalent edge-spread width of $\sigma = 555 \oplus 102~\mu\mathrm{m}$. These results establish a cold-neutron beamline benchmark for a single-converter BGEM detector with full-strip APV25 readout.

physics.ins-det

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safety guarantees transfer to deployment under distribution shift, using diabetes management as a safety-critical testbed. We benchmark safe RL algorithms on a unified clinical simulator and reveal a safety generalization gap: policies satisfying constraints during training frequently violate safety requirements on unseen patients. We demonstrate that test-time shielding, which filters unsafe actions using learned dynamics models, effectively restores safety across algorithms and patient populations. Across eight safe RL algorithms, three diabetes types, and three age groups, shielding achieves Time-in-Range gains of 13--14\% for strong baselines such as PPO-Lag and CPO while reducing clinical risk index and glucose variability. Our simulator and benchmark provide a platform for studying safety under distribution shift in safety-critical control domains. Code is available at https://github.com/safe-autonomy-lab/GlucoSim and https://github.com/safe-autonomy-lab/GlucoAlg.

cs.LG

Safe In-Context Reinforcement Learning

In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history. While ICRL has shown impressive generalization, safety during this adaptation process remains unexplored, limiting its applicability in real-world deployments where test-time behavior is expected to be safe. In this work, we propose SCARED: Safe Contextual Adaptive Reinforcement via Exact-penalty Dual, the first method that promotes safe adaptation of ICRL under the constrained Markov decision process framework. During the parameter-update-free adaptation process, our agent not only maximizes the reward but also keeps the accumulated cost within a user-specified safety budget. We also demonstrate that the agent actively reacts to the safety budget; with a higher safety budget, the agent behaves more aggressively, and with a lower safety budget the agent behaves more conservatively. Across challenging benchmarks, SCARED consistently enables safe and robust in-context adaptation, outperforming existing ICRL and safe meta-RL baselines.

cs.LG

Adaptive Shielding for Safe Reinforcement Learning under Hidden-Parameter Dynamics Shifts

Unseen shifts in environment dynamics, driven by hidden parameters such as friction or gravity, create a challenge for maintaining safety. We address this challenge by proposing Adaptive Shielding, a framework for safe reinforcement learning in constrained hidden-parameter Markov decision processes. A function encoder infers a low-dimensional representation of the underlying dynamics online from transition data, allowing the shield to adapt. To ensure safety during this process, we use a two-layer strategy. First, we introduce safety-regularized optimization that proactively trains the policy away from high-cost regions. Second, the adaptive shielding reactively uses the inferred dynamics to forecast safety risks and applies uncertainty-aware bounds using conformal prediction to filter unsafe actions. We prove that prediction errors in the shielding connect with bounds on the average cost rate. Empirically, across Safe-Gym benchmarks with varying hidden parameters, our approach outperforms baselines on the return-safety trade-off and generalizes reliably to unseen dynamics, while incurring only modest execution-time overhead. Code is available at https://github.com/safe-autonomy-lab/AdaptiveShieldingFE.

cs.LG

Adaptive Reward Design for Reinforcement Learning

There is a surge of interest in using formal languages such as Linear Temporal Logic (LTL) to precisely and succinctly specify complex tasks and derive reward functions for Reinforcement Learning (RL). However, existing methods often assign sparse rewards (e.g., giving a reward of 1 only if a task is completed and 0 otherwise). By providing feedback solely upon task completion, these methods fail to encourage successful subtask completion. This is particularly problematic in environments with inherent uncertainty, where task completion may be unreliable despite progress on intermediate goals. To address this limitation, we propose a suite of reward functions that incentivize an RL agent to complete a task specified by an LTL formula as much as possible, and develop an adaptive reward shaping approach that dynamically updates reward functions during the learning process. Experimental results on a range of benchmark RL environments demonstrate that the proposed approach generally outperforms baselines, achieving earlier convergence to a better policy with higher expected return and task completion rate.

cs.RO

Microscope Projection Photolithography Based on Ultraviolet Light-emitting Diodes

We adapted a conventional microscope for projection photolithography using an ultraviolet (UV) light-emitting diode (LED) as a light source. The use of a UV LED provides the microscope projector with several advantages in terms of compactness and cost. The adapted microscope was capable of producing line patterns as wide as 5 um with the use of a 4X objective lens under optimal lithography conditions. The obtained line width is close to that of the diffraction limit, implying that the line width can be reduced further with the use of a higher resolution photomask and higher magnification objective lens. We expect that low-cost microscope projection photolithography based on a UV LED contribute to the field of physics education or various areas of research, such as chemistry and biology, in the future.

physics.ed-ph