SearcharxivSearch

arXiv subjects

Chengjie Wu

Publications and source records attributed to Chengjie Wu.

9 recordsLinked to original sources

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.

cs.LG

Anomalous current-electric field characteristics in transport through a nanoelectromechanical systems

A deep understanding of the correlation between electronic and mechanical degrees of freedom is crucial to the development of quantum devices in a nanoelectromechanical system (NEMS). In this work, we first establish a fully quantum mechanical approach for transport through a NEMS device, which is valid for arbitrary bias voltages, temperatures, and electro-mechanical couplings. We find an anomalous current-electric field characteristics at a low bias, where the current decreases with a rising electric field, associated with the backward tunneling of electrons for a weak mechanical damping. We reveal that this intriguing behavior arises from a combined effect of mechanical motion and Coulomb blockade, where the rapid increase of backward tunneling events at a large oscillation amplitude suppresses the forward current due to prohibition of double occupation. In the opposite limit of strong damping, the oscillator dissipates its energy to the environment and relaxes to the ground state rapidly. Electrons then transport via the lowest vibrational state such that the net current and its corresponding noise have a vanishing dependence on the electric field.

cond-mat.mes-hall

Fewer May Be Better: Enhancing Offline Reinforcement Learning with Reduced Dataset

Offline reinforcement learning (RL) represents a significant shift in RL research, allowing agents to learn from pre-collected datasets without further interaction with the environment. A key, yet underexplored, challenge in offline RL is selecting an optimal subset of the offline dataset that enhances both algorithm performance and training efficiency. Reducing dataset size can also reveal the minimal data requirements necessary for solving similar problems. In response to this challenge, we introduce ReDOR (Reduced Datasets for Offline RL), a method that frames dataset selection as a gradient approximation optimization problem. We demonstrate that the widely used actor-critic framework in RL can be reformulated as a submodular optimization objective, enabling efficient subset selection. To achieve this, we adapt orthogonal matching pursuit (OMP), incorporating several novel modifications tailored for offline RL. Our experimental results show that the data subsets identified by ReDOR not only boost algorithm performance but also do so with significantly lower computational complexity.

cs.LG

Bayesian Design Principles for Offline-to-Online Reinforcement Learning

Offline reinforcement learning (RL) is crucial for real-world applications where exploration can be costly or unsafe. However, offline learned policies are often suboptimal, and further online fine-tuning is required. In this paper, we tackle the fundamental dilemma of offline-to-online fine-tuning: if the agent remains pessimistic, it may fail to learn a better policy, while if it becomes optimistic directly, performance may suffer from a sudden drop. We show that Bayesian design principles are crucial in solving such a dilemma. Instead of adopting optimistic or pessimistic policies, the agent should act in a way that matches its belief in optimal policies. Such a probability-matching agent can avoid a sudden performance drop while still being guaranteed to find the optimal policy. Based on our theoretical findings, we introduce a novel algorithm that outperforms existing methods on various benchmarks, demonstrating the efficacy of our approach. Overall, the proposed approach provides a new perspective on offline-to-online RL that has the potential to enable more effective learning from offline data.

cs.LG

SEIHAI: A Sample-efficient Hierarchical AI for the MineRL Competition

The MineRL competition is designed for the development of reinforcement learning and imitation learning algorithms that can efficiently leverage human demonstrations to drastically reduce the number of environment interactions needed to solve the complex \emph{ObtainDiamond} task with sparse rewards. To address the challenge, in this paper, we present \textbf{SEIHAI}, a \textbf{S}ample-\textbf{e}ff\textbf{i}cient \textbf{H}ierarchical \textbf{AI}, that fully takes advantage of the human demonstrations and the task structure. Specifically, we split the task into several sequentially dependent subtasks, and train a suitable agent for each subtask using reinforcement learning and imitation learning. We further design a scheduler to select different agents for different subtasks automatically. SEIHAI takes the first place in the preliminary and final of the NeurIPS-2020 MineRL competition.

cs.LG

Celebrating Diversity in Shared Multi-Agent Reinforcement Learning

Recently, deep multi-agent reinforcement learning (MARL) has shown the promise to solve complex cooperative tasks. Its success is partly because of parameter sharing among agents. However, such sharing may lead agents to behave similarly and limit their coordination capacity. In this paper, we aim to introduce diversity in both optimization and representation of shared multi-agent reinforcement learning. Specifically, we propose an information-theoretical regularization to maximize the mutual information between agents' identities and their trajectories, encouraging extensive exploration and diverse individualized behaviors. In representation, we incorporate agent-specific modules in the shared neural network architecture, which are regularized by L1-norm to promote learning sharing among agents while keeping necessary diversity. Empirical results show that our method achieves state-of-the-art performance on Google Research Football and super hard StarCraft II micromanagement tasks.

cs.LG

Towards robust and domain agnostic reinforcement learning competitions

Reinforcement learning competitions have formed the basis for standard research benchmarks, galvanized advances in the state-of-the-art, and shaped the direction of the field. Despite this, a majority of challenges suffer from the same fundamental problems: participant solutions to the posed challenge are usually domain-specific, biased to maximally exploit compute resources, and not guaranteed to be reproducible. In this paper, we present a new framework of competition design that promotes the development of algorithms that overcome these barriers. We propose four central mechanisms for achieving this end: submission retraining, domain randomization, desemantization through domain obfuscation, and the limitation of competition compute and environment-sample budget. To demonstrate the efficacy of this design, we proposed, organized, and ran the MineRL 2020 Competition on Sample-Efficient Reinforcement Learning. In this work, we describe the organizational outcomes of the competition and show that the resulting participant submissions are reproducible, non-specific to the competition environment, and sample/resource efficient, despite the difficult competition task.

cs.LG

Theoretical Exploration on the Magnetic Properties of Ferromagnetic Metallic Glass: An Ising Model on Random Recursive Lattice

The ferromagnetic Ising spins are modeled on a recursive lattice constructed from random-angled rhombus units with stochastic configurations, to study the magnetic properties of the bulk Fe-based metallic glass. The integration of spins on the structural glass model well represents the magnetic moments in the glassy metal. The model is exactly solved by the recursive calculation technique. The magnetization of the amorphous Ising spins, i.e. the glassy metallic magnet is investigated by our modeling and calculation on a theoretical base. The results show that the glassy metallic magnets has a lower Curie temperature, weaker magnetization, and higher entropy comparing to the regular ferromagnet in crystal form. These findings can be understood with the randomness of the amorphous system, and agrees well with others' experimental observations.

cond-mat.mtrl-sci