Searcharxiv⌕ Search

arXiv · 2609.39507

LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation

Abstract

General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot manipulation tasks. Rather than asking agents to submit task-level Python control programs or operate through high-level robot skills, LIBERO-Agent provides an interactive robotic environment where agents can select which observations to inspect, process them with their own tools, and issue native action commands. LIBERO-Agent integrates 200 tasks into a common interaction framework and provides a 30-task primary suite that separates perception, short-horizon execution, and long-horizon composition. Results reveal a pronounced reliability gap: while agents perform well on perception and easy short-horizon tasks, their performance degrades substantially on hard short-horizon and long-horizon tasks. Richer observations improve short-horizon manipulation, while demonstration benefits depend on the agent and format. Among these agents, GPT-6 Astra achieves the strongest overall performance. Further analysis shows its major advantage lies in mechanism interaction, especially when sustained physical contact is needed, while its remaining failures stem from cross-stage interference and geometric errors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zijie Diao, Yitong Chen, Sicheng Xie, Tianyi Lu, Wujian Peng, Guojin Zhong, Houze Xu, Ziyi Ye, Zuxuan Wu, Yu-Gang Jiang. 2026-09-30. LIBERO-Agent: Evaluating General-Purpose Agents for Direct Embodied Manipulation. https://arxiv.org/abs/2609.39507

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer

The sim-to-real gap remains a critical challenge in robotics, hindering the deployment of algorithms trained in simulation to real-world systems. We propose a flexible Real-to-Sim-to-Real (RSR) framework whose central contribution is an information-theoretic cost function that explicitly accounts for sim-to-real discrepancies. This cost balances two objectives, completing the task and steering the policy to collect real-world samples that are maximally informative for improving transfer. It can be integrated seamlessly into existing reinforcement learning algorithms (e.g., PPO, SAC) and ensures a balanced exploration of critical regions in the real domain. The framework treats differentiable simulation as optional: when a differentiable simulator is available, the collected informative data can also be used to tune simulator parameters. We implement the framework with the MuJoCo MJX platform and demonstrate its generality by evaluating on both manipulation tasks with a 6-DOF robotic arm and locomotion tasks on a legged robot. Empirical results show that our RSR loop yields more efficient data acquisition and substantially improves task performance in real-world that achieves a smoother sim-to-real transfer.

cs.RO↗

Learning-Based Progressive Barrier Control for Robot Manipulators with Initial Errors Outside Prescribed Tracking Bounds

Robot manipulators may start a new task with a tracking error larger than the prescribed tolerance. Conventional barrier controllers generally require the initial error to lie within this tolerance, which prevents their direct use under such conditions. This paper develops a progressive barrier controller that gradually contracts an initial error bound to the required value within a prescribed time. The robot can therefore start outside the final bound, while the direct joint-position error satisfies it after the transition. The closed-form control law combines progressive barrier feedback with an online adaptive torque term based on an actor--critic structure. A Lyapunov analysis establishes bounded closed-loop signals and gives sufficient conditions for satisfaction of the final tracking bound. Two-link simulations consider large initial errors, actuator saturation, dynamic variations, disturbances, and measurement errors. The adaptive term reduces the median root-mean-square tracking error by 45.1\% compared with the zero-weight progressive barrier controller. The simulations also map the initial errors that can be handled at different transition times under fixed torque limits. Finally, an experiment on a Niryo Ned3 Pro illustrates tracking performance using measured position and motor-current data.

cs.RO↗

BIM Informed Visual SLAM for Construction Environments

Monitoring building construction sites requires comparing the as-planned design with the as-built state, which can be estimated in real time using Simultaneous Localization and Mapping (SLAM) techniques. However, visual SLAM is prone to trajectory drift in construction environments, producing maps that are geometrically inaccurate with the actual environment. To address this limitation, we augment an existing RGB-D SLAM system with structural priors derived from the Building Information Model (BIM). The system associates detected walls with their BIM counterparts and includes these correspondences as geometric constraints in the back-end optimization, reducing drift and enhancing global consistency. The proposed method operates in real time and is validated on multiple real construction sites, achieving an average trajectory error reduction of 25.23% and a 7.14% improvement in map accuracy over state-of-the-art baselines. Robustness analyses further demonstrate resilience to incomplete BIM data and geometric discrepancies between as-planned models and the as-built environment.

cs.RO↗