Searcharxiv⌕ Search

arXiv subjects

Hikaru Hoshino

Publications and source records attributed to Hikaru Hoshino.

17 recordsLinked to original sources

Learning Stability of Replay-Based Co-Optimization for Transmission Expansion under Strategic Bidding

This paper investigates the behavior of learning-based co-optimization for transmission expansion under strategic bidding in electricity markets. In this framework, transmission capacities are updated while market participants simultaneously learn their bidding strategies through deep reinforcement learning, resulting in coupled and non-stationary learning dynamics. We show that transient policy degradation of bidding agents can generate inconsistent cost-capacity samples, which bias the transmission-capacity update and prevent the co-optimization process from converging to the desired solution. To mitigate this, replay-based capacity updates with nearest-neighbor filtering are introduced to exclude inconsistent samples from the update data. We then analyze a new oscillatory behavior that appears when the replay memory is enlarged. Numerical results on the IEEE 30-bus system demonstrate that enlarged replay memories improve robustness against transient policy degradation but can introduce a temporal lag, leading to oscillations unless the capacity-update learning rate is appropriately reduced. These results reveal the trade-off between replay-memory size and learning rate and provide practical guidelines for stable co-optimization.

eess.SY↗

Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis

Stochastic reach-avoid analysis of controlled dynamical systems is an important tool for safety-critical control under uncertainty, in which the reach-avoid probability is characterized by a Hamilton-Jacobi partial differential equation (PDE). However, solving this PDE using conventional numerical methods becomes computationally intractable as the system dimension increases. Physics-informed neural networks (PINNs) may converge to inaccurate local minima when trained primarily through PDE-residual minimization. Reinforcement learning (RL) offers a scalable alternative, but its learned value functions may be inaccurate or inconsistent with the governing PDE. This paper proposes a physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis. We develop a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function. PDE-residual and boundary-condition losses are then introduced progressively to enforce consistency with the governing PDE and its boundary conditions. The proposed method mitigates the failure modes of conventional PINN techniques while achieving accuracy comparable to that of successfully trained PINNs. The effectiveness of the proposed framework is demonstrated through two case studies.

eess.SY↗

Optimal design of solar-battery hybrid resources considering multi-market participation under weather and price uncertainty

The rapid growth of variable renewable energy has increased the need for flexible and efficiently coordinated energy resources. In this context, hybrid resources that combine renewable generation and battery storage within a single market-participating entity have attracted growing attention. Such hybrid resources can have multiple revenue streams, while allocating limited power and energy capacity across multiple electricity markets including energy and ancillary services. This multi-market coordination increases operational complexity and complicates profitability assessment, making optimal system sizing a challenging design problem. In addition, uncertainty in renewable generation and market prices makes it difficult for conventional optimization approaches to determine system designs that remain effective under stochastic operating conditions. To address these challenges, this paper proposes a deep reinforcement learning-based co-optimization framework for hybrid solar-battery resources. The framework embeds system design variables directly into the policy learning process, enabling joint optimization of hybrid system sizing and coordinated multi-market bidding strategies within a unified stochastic formulation. Case studies using historical renewable generation and market data demonstrate the effectiveness of the proposed framework in identifying economically rational hybrid system design considering multi-market operation.

eess.SY↗

Online Adaptive Probabilistic Safety Certificate with Language Guidance

Achieving long-term safety in uncertain/extreme environments while accounting for human preferences remains a fundamental challenge for autonomous systems. Existing methods often trade off long-term guarantees for fast real-time control and cannot adapt to variability in human preferences or risk tolerance. To address these limitations, we propose a language-guided adaptive probabilistic safety certificate (PSC) framework that guarantees long-term safety for stochastic systems under environmental uncertainty while accommodating diverse human preferences. The proposed framework integrates natural-language inputs from users and Bayesian estimators of the environment into adaptive safety certificates that explicitly account for user preferences, system dynamics, and quantified uncertainties. Our key technical innovation leverages probabilistic invariance--a generalization of forward invariance to a probability space--to obtain myopic safety conditions with long-term safety guarantees. We validate the framework through numerical simulations of autonomous lane-keeping with human-in-the-loop guidance under uncertain and extreme road conditions, demonstrating enhanced safety-performance trade-offs, adaptability to changing environments, and personalization to different user preferences. Code is available at https://github.com/hoshino06/adaptive_lane_keeping.

eess.SY↗

A Reinforcement Learning-based Transmission Expansion Framework Considering Strategic Bidding in Electricity Markets

Transmission expansion planning in electricity markets is tightly coupled with the strategic bidding behaviors of generation companies. This paper proposes a Reinforcement Learning (RL)-based co-optimization framework that simultaneously learns transmission investment decisions and generator bidding strategies within a unified training process. Based on a multiagent RL framework for market simulation, the proposed method newly introduces a design policy layer that jointly optimizes continuous/discrete transmission expansion decisions together with strategic bidding policies. Through iterative interaction between market clearing and investment design, the framework effectively captures their mutual influence and achieves consistent co-optimization of expansion and bidding decisions. Case studies on the IEEE 30-bus system are provided for proof-of-concept validation of the proposed co-optimization framework.

eess.SY↗

Sizing of Battery Considering Renewable Energy Bidding Strategy with Reinforcement Learning

This paper proposes a novel computationally efficient algorithm for optimal sizing of Battery Energy Storage Systems (BESS) considering renewable energy bidding strategies. Unlike existing two-stage methods, our algorithm enables the cooptimization of both by updating the BESS size during the training of the bidding policy, leveraging an extended reinforcement learning (RL) framework inspired by advancements in embodied cognition. By integrating the Deep Recurrent Q-Network (DRQN) with a distributed RL framework, the proposed algorithm effectively manages uncertainties in renewable generation and market prices while enabling parallel computation for efficiently handling long-term data.

eess.SY↗

Probabilistic Reachability Analysis of Multi-scale Voltage Dynamics Using Reinforcement Learning

Voltage stability in modern power systems involves coupled dynamics across multiple time scales. Conventional methods based on time-scale separation or static stability margins may overlook instabilities caused by the coupling of slow and fast transients. Uncertainty in operating conditions further complicates stability assessment, and high computational cost of Monte Carlo simulations limit its applicability to multi-scale dynamics. This paper presents a deep reinforcement learning-based framework for probabilistic reachability analysis of multi-scale voltage dynamics. By formulating each instability mechanism as a distinct absorbing state and introducing a multi-critic architecture for mechanism-specific learning, the proposed method enables consistent learning of risk probabilities associated with multiple instability types within a unified framework. The approach is demonstrated on a four-bus system with load tap changers and over-excitation limiters, illustrating effectiveness of the proposed learning-based reachability analysis in identifying and quantifying the mechanisms leading to voltage collapse.

eess.SY↗

Combined Plant and Control Co-design via Solutions of Hamilton-Jacobi-Bellman Equation Based on Physics-informed Learning

This paper addresses integrated design of engineering systems, where physical structure of the plant and controller design are optimized simultaneously. To cope with uncertainties due to noises acting on the dynamics and modeling errors, an Uncertain Control Co-design (UCCD) problem formulation is proposed. Existing UCCD methods usually rely on uncertainty propagation analyses using Monte Calro methods for open-loop solutions of optimal control, which suffer from stringent trade-offs among accuracy, time horizon, and computational time. The proposed method utilizes closed-loop solutions characterized by the Hamilton-Jacobi-Bellman equation, a Partial Differential Equation (PDE) defined on the state space. A solution algorithm for the proposed UCCD formulation is developed based on PDE solutions of Physics-informed Neural Networks (PINNs). Numerical examples of regulator design problems are provided, and it is shown that simultaneous update of PINN weights and the design parameters effectively works for solving UCCD problems.

eess.SY↗

Autonomous Drifting Based on Maximal Safety Probability Learning

This paper proposes a novel learning-based framework for autonomous driving based on the concept of maximal safety probability. Efficient learning requires rewards that are informative of desirable/undesirable states, but such rewards are challenging to design manually due to the difficulty of differentiating better states among many safe states. On the other hand, learning policies that maximize safety probability does not require laborious reward shaping but is numerically challenging because the algorithms must optimize policies based on binary rewards sparse in time. Here, we show that physics-informed reinforcement learning can efficiently learn this form of maximally safe policy. Unlike existing drift control methods, our approach does not require a specific reference trajectory or complex reward shaping, and can learn safe behaviors only from sparse binary rewards. This is enabled by the use of the physics loss that plays an analogous role to reward shaping. The effectiveness of the proposed approach is demonstrated through lane keeping in a normal cornering scenario and safe drifting in a high-speed racing scenario.

cs.RO↗

Model Predictive Online Trajectory Planning for Adaptive Battery Discharging in Fuel Cell Vehicle

This paper presents an online trajectory planning approach for optimal coordination of Fuel Cell (FC) and battery in plug-in Hybrid Electric Vehicle (HEV). One of the main challenges in energy management of plug-in HEV is generating State-of-Charge (SOC) reference curves by optimally depleting battery under high uncertainties in driving scenarios. Recent studies have begun to explore the potential of utilizing partial trip information for optimal SOC trajectory planning, but dynamic responses of the FC system are not taken into account. On the other hand, research focusing on dynamic operation of FC systems often focuses on air flow management, and battery has been treated only partially. Our aim is to fill this gap by designing an online trajectory planner for dynamic coordination of FC and battery systems that works with a high-level SOC planner in a hierarchical manner. We propose an iterative LQR based online trajectory planning method where the amount of electricity dischargeable at each driving segment can be explicitly and adaptively specified by the high-level planner. Numerical results are provided as a proof of concept example to show the effectiveness of the proposed approach.

eess.SY↗

Physics-informed RL for Maximal Safety Probability Estimation

Accurate risk quantification and reachability analysis are crucial for safe control and learning, but sampling from rare events, risky states, or long-term trajectories can be prohibitively costly. Motivated by this, we study how to estimate the long-term safety probability of maximally safe actions without sufficient coverage of samples from risky states and long-term trajectories. The use of maximal safety probability in control and learning is expected to avoid conservative behaviors due to over-approximation of risk. Here, we first show that long-term safety probability, which is multiplicative in time, can be converted into additive costs and be solved using standard reinforcement learning methods. We then derive this probability as solutions of partial differential equations (PDEs) and propose Physics-Informed Reinforcement Learning (PIRL) algorithm. The proposed method can learn using sparse rewards because the physics constraints help propagate risk information through neighbors. This suggests that, for the purpose of extracting more information for efficient learning, physics constraints can serve as an alternative to reward shaping. The proposed method can also estimate long-term risk using short-term samples and deduce the risk of unsampled states. This feature is in stark contrast with the unconstrained deep RL that demands sufficient data coverage. These merits of the proposed method are demonstrated in numerical simulation.

eess.SY↗

Iterative Linear Quadratic Regulator With Variational Equation-Based Discretization

This paper discusses discretization methods for implementing nonlinear model predictive controllers using Iterative Linear Quadratic Regulator (ILQR). Finite-difference approximations are mostly used to derive a discrete-time state equation from the original continuous-time model. However, the timestep of the discretization is sometimes restricted to be small to suppress the approximation error. In this paper, we propose to use the variational equation for deriving linearizations of the discretized system required in ILQR algorithms, which allows accurate computation regardless of the timestep. Numerical simulations of the swing-up control of an inverted pendulum demonstrate the effectiveness of this method. By the relaxing stringent requirement for the size of the timestep, the use of the variational equation can improve control performance by increasing the number of ILQR iterations possible at each timestep in the realtime computation.

eess.SY↗

Simultaneous Modeling of In Vivo and In Vitro Effects of Nondepolarizing Neuromuscular Blocking Drugs

Nondepolarizing neuromuscular blocking drugs (NDNBs) are clinically used to produce muscle relaxation during general anesthesia. This paper explores a suitable model structure to simultaneously describe in vivo and in vitro effects of three clinically used NDNBs, cisatracurium, vecuronium, and rocuronium. In particular, it is discussed how to reconcile an apparent discrepancy that rocuronium is less potent at inducing muscle relaxation in vivo than predicted from in vitro experiments. We develop a framework for estimating model parameters from published in vivo and in vitro data, and thereby compare the descriptive abilities of several candidate models. It is found that modeling of dynamic effect of activation of acetylcholine receptors (AChRs) is essential for describing in vivo experimental results, and a cyclic gating scheme of AChRs is suggested to be appropriate. Furthermore, it is shown that the above discrepancy in experimental results can be resolved when we consider the fact that the in vivo concentration of ACh is quite low to activate only a part of AChRs, whereas more than 95% of AChRs are activated during in vitro experiments, and that the site-selectivity is smaller for rocuronium than those for cisatracurium and vecuronium.

eess.SY↗

Screening Curve Method for Economic Analysis of Household Solar Energy Self-Consumption

The profitability of solar energy self-consumption in households, the so-called photovoltaic (PV) self-consumption, is expected to boost the deployment of PV and battery storage systems. This paper develops a novel method for economic analysis of PV self-consumption using battery storage based on an extension of the Screening Curve Method (SCM). The SCM enables quick and intuitive estimation of the least-cost generation mix for a target load curve and has been used for generation planning for bulk power systems. In this paper, we generalize the framework of existing SCM to take into account the intermittent nature of renewable energy sources and apply it to the problem of optimal sizing of PV and battery storage systems for a household. Numerical studies are provided to verify the estimation accuracy of the proposed SCM and to illustrate its effectiveness in a sensitivity analysis, owing to its ability to show intuitive plots of cost curves for researchers or policy-makers to understand the reasons behind the optimization results.

eess.SY↗

Model Predictive Control of Smart Districts Participating in Frequency Regulation Market: A Case Study of Using Heating Network Storage

Flexibility provided by Combined Heat and Power (CHP) units in district heating networks is an important means to cope with increasing penetration of intermittent renewable energy resources, and various methods have been proposed to exploit thermal storage tanks installed in these networks. This paper studies a novel problem motivated by an example of district heating and cooling networks in Japan, where high-temperature steam is used as the heating medium. In steam-based networks, storage tanks are usually absent, and there is a strong need to utilize thermal inertia of the pipeline network as storage. However, this type of use of a heating network directly affects the operating condition of the network, and assuring safety and supply quality at the use side is an open problem. To address this, we formulate a novel control problem to utilize CHP units in frequency regulation market while satisfying physical constraints on a steam network described by a nonlinear model capturing dynamics of heat flows and heat accumulation in the network. Furthermore, a Model Predictive Control (MPC) framework is proposed to solve this problem. By consistently combining several nonlinear control techniques, a computationally efficient MPC controller is obtained and shown to work in real-time.

eess.SY↗

A Lumped-Parameter Model of Multiscale Dynamics in Steam Supply Systems

This paper focuses on multiscale dynamics occurring in steam supply systems. The dynamics of interest are originally described by a distributed-parameter model for fast steam flows over a pipe network coupled with a lumped-parameter model for slow internal dynamics of boilers. We derive a lumped-parameter model for the dynamics through physically-relevant approximations. The derived model is then analyzed theoretically and numerically in terms of existence of normally hyperbolic invariant manifold in the phase space of the model. The existence of the manifold is a dynamical evidence that the derived model preserves the slow-fast dynamics, and suggests a separation principle of short-term and long-term operations of steam supply systems, which is analogue to electric power systems. We also quantitatively verify the correctness of the derived model by comparison with brute-force simulation of the original model.

math.DS↗

Structural Analysis and Control of a Model of Two-site Electricity and Heat Supply

This paper introduces a control problem of regulation of energy flows in a two-site electricity and heat supply system, where two Combined Heat and Power (CHP) plants are interconnected via electricity and heat flows. The control problem is motivated by recent development of fast operation of CHP plants to provide ancillary services of power system on the order of tens of seconds to minutes. Due to the physical constraint that the responses of the heat subsystem are not necessary as fast as those of the electric subsystem, the target controlled state is not represented by any isolated equilibrium point, implying that stability of the system is lost in the long-term sense on the order of hours. In this paper, we first prove in the context of nonlinear control theory that the state-space model of the two-site system is non-minimum phase due to nonexistence of isolated equilibrium points of the associated zero dynamics.Instead, we locate a one-dimensional invariant manifold that represents the target controlled flows completely. Then, by utilizing a virtual output under which the state-space model becomes minimum phase, we synthesize a controller that achieves not only the regulation of energy flows in the short-term regime but also stabilization of an equilibrium point in the long-term regime. Effectiveness of the synthesized controller is established with numerical simulations with a practical set of model parameters.

math.OC↗