SearcharxivSearch

arXiv subjects

Yonghyeon Jo

Publications and source records attributed to Yonghyeon Jo.

8 recordsLinked to original sources

Retaining Suboptimal Actions to Follow Shifting Optima in Multi-Agent Reinforcement Learning

Value decomposition is a core approach for cooperative multi-agent reinforcement learning (MARL). However, existing methods still rely on a single optimal action and struggle to adapt when the underlying value function shifts during training, often converging to suboptimal policies. To address this limitation, we propose Successive Sub-value Q-learning (S2Q), which learns multiple sub-value functions to retain alternative high-value actions. Incorporating these sub-value functions into a Softmax-based behavior policy, S2Q encourages persistent exploration and enables $Q^{\text{tot}}$ to adjust quickly to the changing optima. Experiments on challenging MARL benchmarks confirm that S2Q consistently outperforms various MARL algorithms, demonstrating improved adaptability and overall performance. Our code is available at https://github.com/hyeon1996/S2Q.

cs.AI

Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning

Cooperation is central to multi-agent reinforcement learning (MARL), yet learned coordination can be fragile when external perturbations disrupt inter-agent interactions. Prior robust MARL methods have primarily considered value-oriented attacks, leaving a gap in robustness when interaction structures themselves are corrupted. In this paper, we propose an interaction-breaking adversarial learning (IBAL) framework that takes an information-theoretic view to construct attacks that impede coordination by perturbing agents' observations and actions, and trains agents to perform reliably under such disruptions. Empirically, our approach improves robustness over existing robust MARL baselines across diverse attack settings and yields stronger performance even under agent-missing scenarios. Our code is available at https://sunwoolee0504.github.io/IBAL.

cs.LG

Shaping Zero-Shot Coordination via State Blocking

Zero-shot coordination (ZSC) aims to enable agents to cooperate with independently trained partners without prior interaction, a key requirement for real-world multi-agent systems and human-AI collaboration. Existing approaches have largely emphasized increasing partner diversity during training, yet such strategies often fall short of achieving reliable generalization to unseen partners. We introduce State-Blocked Coordination (SBC), a simple yet effective framework that improves ZSC by inducing diverse interaction scenarios without direct environment modification. Specifically, SBC generates a family of virtual environments through state blocking, allowing agents to experience a wide range of suboptimal partner policies. Across multiple benchmarks, SBC demonstrates superior performance in zero-shot coordination, including strong generalization to human partners.

cs.LG

Wolfpack Adversarial Attack for Robust Multi-Agent Reinforcement Learning

Traditional robust methods in multi-agent reinforcement learning (MARL) often struggle against coordinated adversarial attacks in cooperative scenarios. To address this limitation, we propose the Wolfpack Adversarial Attack framework, inspired by wolf hunting strategies, which targets an initial agent and its assisting agents to disrupt cooperation. Additionally, we introduce the Wolfpack-Adversarial Learning for MARL (WALL) framework, which trains robust MARL policies to defend against the proposed Wolfpack attack by fostering systemwide collaboration. Experimental results underscore the devastating impact of the Wolfpack attack and the significant robustness improvements achieved by WALL. Our code is available at https://github.com/sunwoolee0504/WALL.

cs.LG

Exclusively Penalized Q-learning for Offline Reinforcement Learning

Constraint-based offline reinforcement learning (RL) involves policy constraints or imposing penalties on the value function to mitigate overestimation errors caused by distributional shift. This paper focuses on a limitation in existing offline RL methods with penalized value function, indicating the potential for underestimation bias due to unnecessary bias introduced in the value function. To address this concern, we propose Exclusively Penalized Q-learning (EPQ), which reduces estimation bias in the value function by selectively penalizing states that are prone to inducing estimation errors. Numerical results show that our method significantly reduces underestimation bias and improves performance in various offline control tasks compared to other offline RL methods

cs.LG

FoX: Formation-aware exploration in multi-agent reinforcement learning

Recently, deep multi-agent reinforcement learning (MARL) has gained significant popularity due to its success in various cooperative multi-agent tasks. However, exploration still remains a challenging problem in MARL due to the partial observability of the agents and the exploration space that can grow exponentially as the number of agents increases. Firstly, in order to address the scalability issue of the exploration space, we define a formation-based equivalence relation on the exploration space and aim to reduce the search space by exploring only meaningful states in different formations. Then, we propose a novel formation-aware exploration (FoX) framework that encourages partially observable agents to visit the states in diverse formations by guiding them to be well aware of their current formation solely based on their own observations. Numerical results show that the proposed FoX framework significantly outperforms the state-of-the-art MARL algorithms on Google Research Football (GRF) and sparse Starcraft II multi-agent challenge (SMAC) tasks.

cs.LG

Exploiting volumetric wave correlation for enhanced depth imaging in scattering medium

Imaging an object embedded within a scattering medium requires the correction of complex sample-induced wave distortions. Existing approaches have been designed to resolve them by optimizing signal waves recorded in each 2D image. Here, we present a volumetric image reconstruction framework that merges two fundamental degrees of freedom, the wavelength and propagation angles of light waves, based on the object momentum conservation principle. On this basis, we propose methods for exploiting the correlation of signal waves from volumetric images to better cope with multiple scattering. By constructing experimental systems scanning both wavelength and illumination angle of the light source, we demonstrated a 32-fold increase in the use of signal waves compared with that of existing 2D-based approaches and achieved ultrahigh volumetric resolution (lateral resolution: 0.41 um, axial resolution: 0.60 um) even within complex scattering medium owing to the optimal coherent use of the extremely broad spectral bandwidth (225 nm).

physics.optics

Near-field imaging beyond the probe aperture limit

Near-field scanning optical microscopy has been an indispensable tool for designing, characterizing and understanding the functionalities of diverse nanoscale photonic devices. As the advances in fabrication technology have driven the devices smaller and smaller, the demand has grown steadily for improving its resolving power, which is determined mainly by the size of the probe attached to the scanner. The use of a smaller probe has been a straightforward approach to increase the resolving power, but it cannot be made arbitrarily small in practice due to the steep reduction of the collection efficiency. Here, we develop a method to enhance the resolving power of near-field imaging beyond the limit set by the physical size of the probe aperture. The main working principle is to unveil high-order near-field eigenmodes invisible with conventional near-field microscopy. The destructive interference of near-field waves is induced in these high-order eigenmodes by the locally varying phases, which can reveal subaperture-scale fine structural details. To extract these eigenmodes, we construct a self-interference near-field microscopy system and measure a fully phase-referenced far- to near-field transmission matrix (FNTM) composed of near-field amplitude and phase maps recorded for various angles of far-field illumination. By the singular value decomposition of the measured FNTM, we could extract the antisymmetric mode, quadrupole mode, and other higher-order modes hidden under the lowest-order symmetric mode. This enables us to resolve double and triple nano-slots whose gap size (50 nm) is three times smaller than the diameter of the probe aperture (150 nm). The subaperture near-field mode mapping by the FTNM can be potentially combined with various existing near-field imaging modalities and promote their ability to interrogate local near-field optical waves of nanoscale devices.

physics.optics