SearcharxivSearch

arXiv subjects

Yoshihiro Mitsuka

Publications and source records attributed to Yoshihiro Mitsuka.

8 recordsLinked to original sources

TLXML: Task-Level Explanation of Meta-Learning via Influence Functions

Meta-learning enables models to rapidly adapt to new tasks by leveraging prior experience, but its adaptation mechanisms remain opaque, especially regarding how past training tasks influence future predictions. We introduce TLXML (Task-Level eXplanation of Meta-Learning), a novel framework that extends influence functions to meta-learning settings and provides task-level explanations of adaptation and inference. By reformulating influence functions for the bi-level structure of meta-learning, we quantify the contribution of each meta-training task to the adapted model's behaviour. To ensure scalability, we propose a Gauss-Newton-based approximation that significantly reduces computational complexity from $O(pq^2)$ to $O(pq)$, where $p$ and $q$ denote the numbers of model and meta parameters, respectively. Moreover, we propose generalized influence functions defined using pseudo-inverse Hessian, which are applicable even when the loss landscape has flat directions. Results demonstrate that TLXML effectively ranks training tasks by their influence on downstream performance, offering concise, intuitive explanations aligned with user-level abstraction. This work provides a critical step toward interpretable and trustworthy meta-learning systems.

cs.LG

Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms

Industry is moving toward autonomous, network-connected machines that detect and adapt to changing conditions, including hardware faults. Conventional fault-tolerant design duplicates hardware and reroutes control logic; reinforcement learning (RL) offers a learning-based alternative. This paper presents the first systematic comparison of two RL algorithms -- Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) -- for integrating fault tolerance into control. Beyond algorithm choice, we investigate four knowledge-transfer strategies: retaining or discarding model parameters, and retaining or discarding storage contents. Performance is evaluated in two Gymnasium environments: Ant-v5 and FetchReachDense-v3. Results show rapid, fault-specific recovery with clear trade-offs. In Ant-v5, retaining PPO's parameters boosts early returns and remains the safest choice across all faults, while retaining SAC's parameters yields mixed outcomes. SAC's early performance further depends on whether the replay buffer is retained: beneficial when prior experiences match current dynamics, but harmful when they diverge. In FetchReachDense-v3, discarding both PPO's and SAC's parameters was most effective under sensor corruption. Across tasks, both algorithms recover near-normal performance within minutes in low-dimensional settings and within days in high-dimensional settings, highlighting a clear trade-off between adaptation speed and asymptotic performance. These findings demonstrate that RL can deliver robust fault tolerance and offer practical guidelines.

cs.LG

WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation

Generative Flow Networks for continuous scenarios (CFlowNets) have shown promise in solving sequential decision-making tasks by learning stochastic policies using a flow and a retrieval network. Despite their demonstrated efficiency compared to state-of-the-art Reinforcement Learning (RL) algorithms, their practical application in robotic control tasks is constrained by the reliance on pre-training the retrieval network. This dependency poses challenges in dynamic robotic environments, where pre-training data may not be readily available or representative of the current environment. This paper introduces WINFlowNets, a novel CFlowNets framework that enables the co-training of flow and retrieval networks. WINFlowNets begins with a warm-up phase for the retrieval network to bootstrap its policy, followed by a shared training architecture and a shared replay buffer for co-training both networks. Experiments in simulated robotic environments demonstrate that WINFlowNets surpasses CFlowNets and state-of-the-art RL algorithms in terms of average reward and training stability. Furthermore, WINFlowNets exhibits strong adaptive capability in fault environments, making it suitable for tasks that demand quick adaptation with limited sample data. These findings highlight WINFlowNets' potential for deployment in dynamic and malfunction-prone robotic systems, where traditional pre-training or sample inefficient data collection may be impractical.

cs.LG

A Study of the Efficacy of Generative Flow Networks for Robotics and Machine Fault-Adaptation

Advancements in robotics have opened possibilities to automate tasks in various fields such as manufacturing, emergency response and healthcare. However, a significant challenge that prevents robots from operating in real-world environments effectively is out-of-distribution (OOD) situations, wherein robots encounter unforseen situations. One major OOD situations is when robots encounter faults, making fault adaptation essential for real-world operation for robots. Current state-of-the-art reinforcement learning algorithms show promising results but suffer from sample inefficiency, leading to low adaptation speed due to their limited ability to generalize to OOD situations. Our research is a step towards adding hardware fault tolerance and fast fault adaptability to machines. In this research, our primary focus is to investigate the efficacy of generative flow networks in robotic environments, particularly in the domain of machine fault adaptation. We simulated a robotic environment called Reacher in our experiments. We modify this environment to introduce four distinct fault environments that replicate real-world machines/robot malfunctions. The empirical evaluation of this research indicates that continuous generative flow networks (CFlowNets) indeed have the capability to add adaptive behaviors in machines under adversarial conditions. Furthermore, the comparative analysis of CFlowNets with reinforcement learning algorithms also provides some key insights into the performance in terms of adaptation speed and sample efficiency. Additionally, a separate study investigates the implications of transferring knowledge from pre-fault task to post-fault environments. Our experiments confirm that CFlowNets has the potential to be deployed in a real-world machine and it can demonstrate adaptability in case of malfunctions to maintain functionality.

cs.RO

Recurrence relations of Kummer functions and Regge string scattering amplitudes

We discover an infinite number of recurrence relations among Regge string scattering amplitudes \cite{bosonic,RRsusy} of different string states at arbitrary mass levels in the open bosonic string theory. As a result, all Regge string scattering amplitudes can be algebraically solved up to multiplicative factors. Instead of decoupling zero-norm states in the fixed angle regime, the calculation is based on recurrence relations and addition theorem of Kummer functions of the second kind. These recurrence relations among Regge string scattering amplitudes are dual to linear relations or symmetries among high-energy fixed angle string scattering amplitudes discovered previously.

hep-th

No more CKY two-forms in the NHEK

We show that in the near-horizon limit of a Kerr-NUT-AdS black hole, the space of conformal Killing-Yano two-forms does not enhance and remains of dimension two. The same holds for an analogous polar limit in the case of extremal NUT charge. We also derive the conformal Killing-Yano $p$-form equation for any background in arbitrary dimension in the form of parallel transport.

gr-qc

Higher Spin String States Scattered from D-particle in the Regge Regime and Factorized Ratios of Fixed Angle Scatterings

We study scattering of higher spin closed string states at arbitrary mass levels from D-particle in the Regge regime. We extract the infinite ratios among high-energy amplitudes of different string states in the fixed angle regime from these Regge string scattering amplitudes. In this calculation, we have used an identity proved recently based on a signless Stirling number identity in combinatorial theory. The complete ratios calculated by this indirect method include a subset of ratios calculated previously by direct fixed angle calculation. Moreover, we discover that in spite of the non-factorizability of the closed string D-particle scattering amplitudes, the complete ratios derived for the fixed angle regime are found to be factorized. These ratios are consistent with the decoupling of high-energy zero norm states calculated previously.

hep-th

Liouville Equation in 1/8 BPS Geometries

We investigate the 1/8 BPS geometries with SU(2) x U(1) x SO(4) x R symmetry in IIB supergravity which were classified by Gava et al, (hep-th/0611065). It is desirable to have a complete set of differential equations imposed on the controlling functions such that they are not only necessary but also sufficient to produce supergravity solutions with those symmetries. We work on this issue and find a new differential equation for the controlling functions. For a special case, we exhaust all the remaining constraints and show that they reduce to one Liouville equation. The solutions of this equation produce geometries which are locally equivalent to the near horizon geometries of intersecting D3-branes.

hep-th