Searcharxiv⌕ Search

arXiv subjects

Kevin Lu

Publications and source records attributed to Kevin Lu.

27 records · Page 2Linked to original sources

Decision Transformer: Reinforcement Learning via Sequence Modeling

We introduce a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture, and associated advances in language modeling such as GPT-x and BERT. In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired return (reward), past states, and actions, our Decision Transformer model can generate future actions that achieve the desired return. Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.

cs.LG↗

Reset-Free Lifelong Learning with Skill-Space Planning

The objective of lifelong reinforcement learning (RL) is to optimize agents which can continuously adapt and interact in changing environments. However, current RL approaches fail drastically when environments are non-stationary and interactions are non-episodic. We propose Lifelong Skill Planning (LiSP), an algorithmic framework for non-episodic lifelong RL based on planning in an abstract space of higher-order skills. We learn the skills in an unsupervised manner using intrinsic rewards and plan over the learned skills using a learned dynamics model. Moreover, our framework permits skill discovery even from offline data, thereby reducing the need for excessive real-world interactions. We demonstrate empirically that LiSP successfully enables long-horizon planning and learns agents that can avoid catastrophic failures even in challenging non-stationary and non-episodic environments derived from gridworld and MuJoCo benchmarks.

cs.LG↗

Efficient Empowerment Estimation for Unsupervised Stabilization

Intrinsically motivated artificial agents learn advantageous behavior without externally-provided rewards. Previously, it was shown that maximizing mutual information between agent actuators and future states, known as the empowerment principle, enables unsupervised stabilization of dynamical systems at upright positions, which is a prototypical intrinsically motivated behavior for upright standing and walking. This follows from the coincidence between the objective of stabilization and the objective of empowerment. Unfortunately, sample-based estimation of this kind of mutual information is challenging. Recently, various variational lower bounds (VLBs) on empowerment have been proposed as solutions; however, they are often biased, unstable in training, and have high sample complexity. In this work, we propose an alternative solution based on a trainable representation of a dynamical system as a Gaussian channel, which allows us to efficiently calculate an unbiased estimator of empowerment by convex optimization. We demonstrate our solution for sample-based unsupervised stabilization on different dynamical control systems and show the advantages of our method by comparing it to the existing VLB approaches. Specifically, we show that our method has a lower sample complexity, is more stable in training, possesses the essential properties of the empowerment function, and allows estimation of empowerment from images. Consequently, our method opens a path to wider and easier adoption of empowerment for various applications.

cs.LG↗

Distinct Distances with $\ell_p$ Spaces

We study Erd\H os's distinct distances problem under $\ell_p$ metrics with integer $p$. We improve the current best bound for this problem from $Ω(n^{4/5})$ to $Ω(n^{6/7-ε})$, for any $ε>0$. We also characterize the sets that span an asymptotically minimal number of distinct distances under the $\ell_1$ and $\ell_\infty$ metrics.

math.CO↗

Weakly supervised one-stage vision and language disease detection using large scale pneumonia and pneumothorax studies

Detecting clinically relevant objects in medical images is a challenge despite large datasets due to the lack of detailed labels. To address the label issue, we utilize the scene-level labels with a detection architecture that incorporates natural language information. We present a challenging new set of radiologist paired bounding box and natural language annotations on the publicly available MIMIC-CXR dataset especially focussed on pneumonia and pneumothorax. Along with the dataset, we present a joint vision language weakly supervised transformer layer-selected one-stage dual head detection architecture (LITERATI) alongside strong baseline comparisons with class activation mapping (CAM), gradient CAM, and relevant implementations on the NIH ChestXray-14 and MIMIC-CXR dataset. Borrowing from advances in vision language architectures, the LITERATI method demonstrates joint image and referring expression (objects localized in the image using natural language) input for detection that scales in a purely weakly supervised fashion. The architectural modifications address three obstacles -- implementing a supervised vision and language detection method in a weakly supervised fashion, incorporating clinical referring expression natural language information, and generating high fidelity detections with map probabilities. Nevertheless, the challenging clinical nature of the radiologist annotations including subtle references, multi-instance specifications, and relatively verbose underlying medical reports, ensures the vision language detection task at scale remains stimulating for future investigation.

cs.CV↗

Adaptive Online Planning for Continual Lifelong Learning

We study learning control in an online reset-free lifelong learning scenario, where mistakes can compound catastrophically into the future and the underlying dynamics of the environment may change. Traditional model-free policy learning methods have achieved successes in difficult tasks due to their broad flexibility, but struggle in this setting, as they can activate failure modes early in their lifetimes which are difficult to recover from and face performance degradation as dynamics change. On the other hand, model-based planning methods learn and adapt quickly, but require prohibitive levels of computational resources. We present a new algorithm, Adaptive Online Planning (AOP), that achieves strong performance in this setting by combining model-based planning with model-free learning. By approximating the uncertainty of the model-free components and the planner performance, AOP is able to call upon more extensive planning only when necessary, leading to reduced computation times, while still gracefully adapting behaviors in the face of unpredictable changes in the world -- even when traditional RL fails.

cs.LG↗

Hamiltonicity in Semi-Regular Tessellation Dual Graphs

This paper shows NP-completeness for finding Hamiltonian cycles in induced subgraphs of the dual graphs of semi-regular tessilations. It also shows NP-hardness for a new, wide class of graphs called augmented square grids. This work follows up on prior studies of the complexity of finding Hamiltonian cycles in regular and semi-regular grid graphs.

cs.CC↗

Weak Subordination of Multivariate Lévy Processes and Variance Generalised Gamma Convolutions

Subordinating a multivariate Lévy process, the subordinate, with a univariate subordinator gives rise to a pathwise construction of a new Lévy process, provided the subordinator and the subordinate are independent processes. The variance-gamma model in finance was generated accordingly from a Brownian motion and a gamma process. Alternatively, multivariate subordination can be used to create Lévy processes, but this requires the subordinate to have independent components. In this paper, we show that there exists another operation acting on pairs $(T,X)$ of Lévy processes which creates a Lévy process $X\odot T$. Here, $T$ is a subordinator, but $X$ is an arbitrary Lévy process with possibly dependent components. We show that this method is an extension of both univariate and multivariate subordination and provide two applications. We illustrate our methods giving a weak formulation of the variance-$α$-gamma process that exhibits a wider range of dependence than using traditional subordination. Also, the variance generalised gamma convolution class of Lévy processes formed by subordinating Brownian motion with Thorin subordinators is further extended using weak subordination.

math.PR↗

A thermodynamic unification of jamming

Fragile materials ranging from sand to fire-retardant to toothpaste are able to exhibit both solid and fluid-like properties across the jamming transition. Unlike ordinary fusion, systems of grains, foams and colloids jam and cease to flow under conditions that still remain unknown. Here we quantify jamming via a thermodynamic approach by accounting for the structural ageing and the shear-induced compressibility of dry sand. Specifically, the jamming threshold is defined using a non-thermal temperature that measures the 'fluffiness' of a granular mixture. The thermodynamic model, casted in terms of pressure, temperature and free-volume, also successfully predicts the entropic data of five molecular glasses. Notably, the predicted configurational entropy avoids the Kauzmann paradox entirely. Without any free parameters, the proposed equation-of-state also governs the mechanism of shear-banding and the associated features of shear-softening and thickness-invariance.

cond-mat.soft↗