SearcharxivSearch

arXiv subjects

Benjamin Poole

Publications and source records attributed to Benjamin Poole.

7 recordsLinked to original sources

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While human demonstrations and feedback have proven crucial for alignment, existing approaches predominantly combine these signals using multi-stage pipelines designed for the contextual bandit framing of language generation. Yet little work explores how these complementary inputs can serve as a richer, interconnected signal for single-stage offline training in fully sequential decision-making environments. We propose Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that harnesses evaluative feedback as a corrective signal to improve the alignment of imitation learning policies. We adapt Safety Gymnasium environments to be a principled testbed for alignment evaluation, demonstrating improved aptitude and up to a 98\% reduction in misalignment across a range of imitation learning algorithms. FMR remains robust in limited data regimes, even when learning from scarce aligned and uninformative noisy demonstrations.

cs.AI

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning

Data rehearsal has emerged as a leading approach for mitigating catastrophic forgetting in Continual Reinforcement Learning (CRL). However, existing work remains confined to policy gradient frameworks, regularizing only actors due to the performance degradation incurred by critic regularization. This actor-centric approach overlooks the potential of data rehearsal for value function approximation. Moreover, existing evaluations in CRL rarely consider multi-cyclic environments where task sequences repeat, a critical real-world scenario that exacerbates forgetting and plasticity. We investigate data rehearsal for Deep Q-Networks using Q-value regularization in multi-cyclic settings and propose Qreg+NWLU which introduces two simple modifications: (1) continuous data rehearsal that dynamically collects and updates stored Q-values throughout training, and (2) "No-Wait" regularization that applies immediately rather than after the first task. Together, these modifications yield improvements in learning efficiency, forgetting mitigation, and knowledge transfer over Qreg and conventional CRL methods within value function approximation settings.

cs.LG

Anomalous, pre-yield grain-boundary sliding in copper revealed with in-situ high-resolution strain mapping

Grain boundary sliding is typically associated with high temperature deformation in engineering alloys. Here, we examine grain boundary sliding at room temperature in oxygen-free high-conductivity copper under quasi-static tensile testing. By using high-resolution digital image correlation (HRDIC) conducted in-situ within a scanning electron microscope to produce time-series strain maps, we unexpectedly observe that grain boundary sliding occurs extensively prior to macroscopic yield, and before the onset of significant crystallographic slip. Extreme values in strain and in-plane rotation are found to be associated with grain boundaries immediately prior to yield and during the initial stages of plastic deformation, which are higher than those associated with crystallographic slip. By combining laser scanning confocal microscopy height mapping with the strain maps and orientation maps from electron backscatter diffraction, grain boundary sliding character is determined, finding evidence of pure in-plane, pure out-of-plane and mixed-mode sliding.

cond-mat.mtrl-sci

Error-related Potential Variability: Exploring the Effects on Classification and Transferability

Brain-Computer Interfaces (BCI) have allowed for direct communication from the brain to external applications for the automatic detection of cognitive processes such as error recognition. Error-related potentials (ErrPs) are a particular brain signal elicited when one commits or observes an erroneous event. However, due to the noisy properties of the brain and recording devices, ErrPs vary from instance to instance as they are combined with an assortment of other brain signals, biological noise, and external noise, making the classification of ErrPs a non-trivial problem. Recent works have revealed particular cognitive processes such as awareness, embodiment, and predictability that contribute to ErrP variations. In this paper, we explore the performance of classifier transferability when trained on different ErrP variation datasets generated by varying the levels of awareness and embodiment for a given task. In particular, we look at transference between observational and interactive ErrP categories when elicited by similar and differing tasks. Our empirical results provide an exploratory analysis into the ErrP transferability problem from a data perspective.

cs.HC

Towards Interactive Reinforcement Learning with Intrinsic Feedback

Reinforcement learning (RL) and brain-computer interfaces (BCI) have experienced significant growth over the past decade. With rising interest in human-in-the-loop (HITL), incorporating human input with RL algorithms has given rise to the sub-field of interactive RL. Adjacently, the field of BCI has long been interested in extracting informative brain signals from neural activity for use in human-computer interactions. A key link between these fields lies in the interpretation of neural activity as feedback such that interactive RL approaches can be employed. We denote this new and emerging medium of feedback as intrinsic feedback. Despite intrinsic feedback's ability to be conveyed automatically and even unconsciously, proper exploration surrounding this key link has largely gone unaddressed by both communities. Thus, to help facilitate a deeper understanding and a more effective utilization, we provide a tutorial-style review covering the motivations, approaches, and open problems of intrinsic feedback and its foundational concepts.

cs.AI

Slip band interactions and GND latent hardening in a galling resistant stainless steel

Slip activation, slip band interactions, and GND densities in iron-base, galling resistant alloy Nitronic 60 have been characterised at the grain length scale using small-scale mechanical testing with high resolution digital image correlation and high-angular resolution electron backscatter diffraction. By correlating the two measurement techniques, new insight into slip band interactions, the generation of lattice curvature and the corresponding accumulation of geometrically necessary dislocations (GNDs) is provided. Multiple discrete slip bands are typically active within single grains, resulting in significant slip band interactions. Crossing slip bands were found to generate accumulations of GNDs. Regions where slip bands block other slip bands were associated with the highest GND densities, in excess of three time the densities of crossing slip bands. Representative crystal plasticity modelling investigations have demonstrated that discrete slip blocking events are responsible for locally elevated GND density. This behaviour is rationalised in terms of lattice curvature associated with the differing levels of constraint provided by the crossing or blocking-type behaviours. Ferrite grains are also found to contribute to the generation of GNDs. Together, these two effects provide significant work hardening mechanisms, likely to be key to the development of future iron-base hard facing alloys.

cond-mat.mtrl-sci

The roles of adhesion, internal heat generation and elevated temperatures in normally loaded, sliding rough surfaces

The thermal effects of plastic and frictional heat generation and elevated temperature were examined along with the role of adhesion in the context of galling wear, using a representative crystal plasticity, normally loaded, sliding surface model. Galling frequency behaviour was predicted for 316L steel. Deformation of the surfaces was dominated by the surface geometry, with no significant effect due to variations in frictional models. Plastic and frictional heating were found to have a minimal effect on the deformation of the surface, with the rapid conduction of heat preventing any highly localised heating. There was no corresponding effect on the predicted galling frequency response. Isothermal, elevated temperature conditions caused a decrease in galling resistance, driven by the temperature sensitivity of the critical resolved shear stress. The extent of deformation, as quantified by the area of plastically deformed material and plastic reach, increased with temperature. Comparisons were made with literature results for several surface amplitude and wavelength conditions. Model results compared favourably with those in the literature. However, the reduction in predicted galling resistance with elevated temperature for a fixed surface was not as severe as observations in the literature, suggesting other mechanisms (e.g. phase transformations, surface coatings and oxides) are likely important.

cond-mat.mtrl-sci