SearcharxivSearch

arXiv subjects

Yawei Wang

Publications and source records attributed to Yawei Wang.

23 records · Page 2Linked to original sources

Reward function shape exploration in adversarial imitation learning: an empirical study

For adversarial imitation learning algorithms (AILs), no true rewards are obtained from the environment for learning the strategy. However, the pseudo rewards based on the output of the discriminator are still required. Given the implicit reward bias problem in AILs, we design several representative reward function shapes and compare their performances by large-scale experiments. To ensure our results' reliability, we conduct the experiments on a series of Mujoco and Box2D continuous control tasks based on four different AILs. Besides, we also compare the performance of various reward function shapes using varying numbers of expert trajectories. The empirical results reveal that the positive logarithmic reward function works well in typical continuous control tasks. In contrast, the so-called unbiased reward function is limited to specific kinds of tasks. Furthermore, several designed reward functions perform excellently in these environments as well.

cs.LG

Autonomous Charging of Electric Vehicle Fleets to Enhance Renewable Generation Dispatchability

A total 19% of generation capacity in California is offered by PV units and over some months, more than 10% of this energy is curtailed. In this research, a novel approach to reduce renewable generation curtailments and increasing system flexibility by means of electric vehicles' charging coordination is represented. The presented problem is a sequential decision making process, and is solved by fitted Q-iteration algorithm which unlike other reinforcement learning methods, needs fewer episodes of learning. Three case studies are presented to validate the effectiveness of the proposed approach. These cases include aggregator load following, ramp service and utilization of non-deterministic PV generation. The results suggest that through this framework, EVs successfully learn how to adjust their charging schedule in stochastic scenarios where their trip times, as well as solar power generation are unknown beforehand.

eess.SY

Wasserstein Distance guided Adversarial Imitation Learning with Reward Shape Exploration

The generative adversarial imitation learning (GAIL) has provided an adversarial learning framework for imitating expert policy from demonstrations in high-dimensional continuous tasks. However, almost all GAIL and its extensions only design a kind of reward function of logarithmic form in the adversarial training strategy with the Jensen-Shannon (JS) divergence for all complex environments. The fixed logarithmic type of reward function may be difficult to solve all complex tasks, and the vanishing gradients problem caused by the JS divergence will harm the adversarial learning process. In this paper, we propose a new algorithm named Wasserstein Distance guided Adversarial Imitation Learning (WDAIL) for promoting the performance of imitation learning (IL). There are three improvements in our method: (a) introducing the Wasserstein distance to obtain more appropriate measure in the adversarial training process, (b) using proximal policy optimization (PPO) in the reinforcement learning stage which is much simpler to implement and makes the algorithm more efficient, and (c) exploring different reward function shapes to suit different tasks for improving the performance. The experiment results show that the learning procedure remains remarkably stable, and achieves significant performance in the complex continuous control tasks of MuJoCo.

cs.LG

Fully superconducting machine for electric aircraft propulsion: study of AC loss for HTS stator

Fully superconducting machines provide the high power density required for future electric aircraft propulsion. However, superconducting windings generate AC losses in AC electrical machine environments. These AC losses are difficult to remove at low temperatures and they add an extra burden to the aircraft cooling system. Due to heavy cooling penalty, AC losses in the HTS stator, is one of the key topics in HTS machine design. In order to evaluate the AC loss of superconducting stator windings in a rotational machine environment, we designed and built a novel axial-flux high temperature superconducting (HTS) machine platform. The AC loss measurement is based on calorimetrically boiling-off liquid nitrogen. Both total AC loss and magnetisation loss in HTS stator are measured in a rotational magnetic field condition. This platform is essential to study ways to minimise AC losses in HTS stator, in order to maximum the efficiency of fully HTS machines.

physics.app-ph

3D quench modeling based on T-A formulation for high temperature superconductor CORC cables

High temperature superconductor (HTS) (RE)Ba2Cu3Ox (REBCO) conductor on round core cable (CORC) has high current carrying capacity for high field magnet and power applications. In REBCO CORC cables, current redistribution occurs among tapes through terminal contact resistances when a local quench occurs. Therefore, the quench behaviour of CORC cable is different from single tape situation, for it is significantly affected by terminal contact resistances. To better understand the underlying physical process of local quenches in CORC cables, a new 3D multi-physics modelling tool for CORC cables is developed and presented in this paper. In this model, the REBCO tape is treated as a thin shell without thickness, and four models are coupled: T-formulation model, A-formulation model, a heat transfer model and an equivalent circuit model. The current redistribution, temperature and tape voltage of CORC cable during hot spot induced quenches are analysed using this model. The results show that the thermal stability of CORC cable can be considerably improved by reducing terminal contact resistance. The minimum quench energy (MQE) increases rapidly with the reduction of terminal contact resistance when the resistance is in a middle range. When the terminal contact resistance is too low or too high, the MQE shows no obvious variation with terminal contact resistances. With a low terminal contact resistance, a hot spot in one tape may induce an over-current quench on the other tapes without hot spots. This will not happen in a cable with high terminal contact resistance. In this case, the tape with hot spot will quench and burn out before inducing a quench on other tapes. The modelling tool developed can be used to design CORC cables with improved thermal stability.

cond-mat.supr-con