Searcharxiv⌕ Search

arXiv · 2610.09891

A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration

Abstract

Human-Robot Collaboration (HRC) can facilitate mass customisation in Industry 4.0, with Reinforcement Learning from Human Feedback (RLHF) representing a promising approach for developing safe AI-based robots. Practical challenges remain regarding safety during AI development, human feedback quality, and bidirectional human-robot adaptation. We conducted a scoping review of RLHF in HRC systems, mapping methods that address these challenges. Following PRISMA guidelines, we screened 199 records and included 20 peer-reviewed publications (2020-2025) spanning multiple HRC domains. To our knowledge, this is the first review focused on the bidirectional, closed-loop design of RLHF. Our review found multiple feedback modalities enabling data collection in various feedback formats. Collected data can be integrated at different stages of AI training, resulting in a multi-step development process. Pilot experiments are commonly used to evaluate HRC systems based on both human and robot metrics. To empirically test a key gap identified in the review, we conducted a between-subjects VR experiment comparing system- and user-initiated feedback on robot proxemic behaviour for safe navigation. Using Bayesian models, we analysed the relation between the collected feedback and safety metrics: psychological safety (post-experiment questionnaire) and physical safety (inverse time-to-collision). Results show that user-initiated feedback captures perceived safety better than system-initiated feedback, indicating that feedback timing directly affects feedback quality. Our review and experiment findings show that RLHF relies on appropriate feedback methods to ensure AI safety in HRC, and future RLHF research should prioritise realistic HRC experiments evaluating the effects of feedback collection methods on relevant human and robot metrics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alexandra Coroiu, Andrea Vogt, Viktor Werbilo, Andreas Poppele, Johann Christensen, Sven Hallerbach. 2026-10-07. A Scoping Review and Experimental Study on Reinforcement Learning from Human Feedback for Human-Robot Collaboration. https://arxiv.org/abs/2610.09891

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Before Bringing It Up: When and How AI Companions Should Use Memory

Memory can sustain AI companionship, yet even accurate recollection can be inappropriate to use. Two rounds of formative interviews with 14 users (n = 6 exploratory, n = 8 memory-focused) motivate asking what a companion should consider before using past information. Eight themes inform Reconsider, a single-call procedure with five checks and four handling modes, evaluated on 80 scenarios across five models over 400 blinded within-model pairs. Two LLM judges favored Reconsider by net margins of +15 and +23 percentage points, with bootstrap intervals excluding zero for three of five models but not for GPT or Claude. Evaluator analysis linked judge scoring differences to model family, and a preliminary matched-guidance control isolating memory-specific content gave positive margins. We contribute an interview-grounded design framework for memory use and an evaluation that scrutinizes its own evaluators.

cs.HC↗

Intent Graph: Navigating the Analytical Reasoning Space for Exploratory Data Analysis

Exploratory data analysis (EDA) is rarely open-ended in practice: analysts work from high-level domain questions toward the concrete analyses that can answer them, prioritizing directions with domain knowledge and prior hypotheses. Large language models (LLMs) can supply such knowledge, but their responses are unstructured, leaving analysts no way to see what has been explored, what is missing, or why one direction was chosen over another. We present DAG-EDA, a system that lets analysts and an LLM co-navigate the space of possible analyses through two linked structures. An intent graph, governed by a grammar of analytical intent, decomposes an ambiguous natural-language question into progressively concrete analysis tasks, keeping alternative framings open and letting analysts branch, backtrack, and compare paths. A multi-layered knowledge graph externalizes the LLM's domain knowledge, linking domain concepts to the dataset variables that can measure them, so analysts can inspect and contest how their question is grounded in the data. Both graphs are constructed from only the dataset and the analyst's question, and the analyses the analyst reaches are rendered as interactive dashboards. We illustrate the system through a usage scenario and describe a user study design for examining whether the system scaffold analysts' reasoning and navigation.

cs.HC↗

Magic Pen: Automatic Pen Mode Switching for Document Annotation

Traditional digital pen interfaces use menu buttons to change the pen mode, which results in time and cognitive load spent on round-trip interactions and mode errors from tapping small mode selection buttons. This work presents the Magic Pen, a technique which uses machine learning to automatically switch between digital pen modes without requiring explicit mode changes. Magic Pen is driven by an LSTM model trained on pen data collected from 27 participants across two studies and uses transfer learning to iteratively tune the model towards how a specific user annotates. Error mitigation techniques using a flick gesture or on-screen tap are incorporated to correct mode errors or remove a stroke quickly. We evaluated Magic Pen in a comparative study with 18 participants, followed by iterative improvements and a deployment study with 8 participants. Magic Pen was preferred compared to a conventional menu-based approach, and transfer learning allowed for greater model predictability and stability.

cs.HC↗