SearcharxivSearch

arXiv subjects

Haibo Zhao

Publications and source records attributed to Haibo Zhao.

11 recordsLinked to original sources

Residual Rotation Correction using Tactile Equivariance

Visuotactile policy learning augments vision-only policies with tactile input, facilitating contact-rich manipulation. However, the high cost of tactile data collection makes sample efficiency the key requirement for developing visuotactile policies. We present EquiTac, a framework that exploits the inherent SO(2) symmetry of in-hand object rotation to improve sample efficiency and generalization for visuotactile policy learning. EquiTac first reconstructs surface normals from raw RGB inputs of vision-based tactile sensors, so rotations of the normal vector field correspond to in-hand object rotations. An SO(2)-equivariant network then predicts a residual rotation action that augments a base visuomotor policy at test time, enabling real-time rotation correction without additional reorientation demonstrations. On a real robot, EquiTac accurately achieves robust zero-shot generalization to unseen in-hand orientations with very few training samples, where baselines fail even with more training data. To our knowledge, this is the first tactile learning method to explicitly encode tactile equivariance for policy learning, yielding a lightweight, symmetry-aware module that improves reliability in contact-rich tasks.

cs.RO

Generalizable Hierarchical Skill Learning via Object-Centric Representation

We present Generalizable Hierarchical Skill Learning (GSL), a novel framework for hierarchical policy learning that significantly improves policy generalization and sample efficiency in robot manipulation. One core idea of GSL is to use object-centric skills as an interface that bridges the high-level vision-language model and the low-level visual-motor policy. Specifically, GSL decomposes demonstrations into transferable and object-canonicalized skill primitives using foundation models, ensuring efficient low-level skill learning in the object frame. At test time, the skill-object pairs predicted by the high-level agent are fed to the low-level module, where the inferred canonical actions are mapped back to the world frame for execution. This structured yet flexible design leads to substantial improvements in sample efficiency and generalization of our method across unseen spatial arrangements, object appearances, and task compositions. In simulation, GSL trained with only 3 demonstrations per task outperforms baselines trained with 30 times more data by 15.5 percent on unseen tasks. In real-world experiments, GSL also surpasses the baseline trained with 10 times more data.

cs.RO

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly focus on task-level evaluation and fail to provide actionable insights into the underlying causes of model failures. To address this limitation, we introduce BEAR, a benchmark that decomposes embodied tasks into 14 atomic skills for fine-grained skill-level evaluation. BEAR comprises 4,469 interleaved image-video-text samples spanning 14 skills across 6 categories, ranging from low-level perception to high-level planning. We evaluate 20 MLLMs on BEAR under a hierarchical skill-level diagnosis framework and uncover two key findings: (1) perceptual capabilities are major bottlenecks behind reasoning failures, and (2) current models suffer from unstable spatiotemporal modeling that remains largely unexposed in prior benchmarks. Motivated by these findings, we further propose BEAR-Agent, a multimodal conversational agent that augments MLLMs with visual and spatial reasoning tools. BEAR-Agent substantially improves performance across embodied skills, achieving a relative improvement of 17.5% on GPT-5 over the base model on BEAR, while also outperforming strong baselines in both simulation and real-world robotic experiments. Project page: https://bear-official66.github.io/

cs.CV

CO-OPERA: A Human-AI Collaborative Playwriting Tool to Support Creative Storytelling for Interdisciplinary Drama Education

Drama-in-education is an interdisciplinary instructional approach that integrates subjects such as language, history, and psychology. Its core component is playwriting. Based on need-finding interviews of 13 teachers, we found that current general-purpose AI tools cannot effectively assist teachers and students during playwriting. Therefore, we propose CO-OPERA - a collaborative playwriting tool integrating generative artificial intelligence capabilities. In CO-OPERA, users can both expand their thinking through discussions with a tutor and converge their thinking by operating agents to generate script elements. Additionally, the system allows for iterative modifications and regenerations based on user requirements. A system usability test conducted with middle school students shows that our CO-OPERA helps users focus on whole logical narrative development during playwriting. Our playwriting examples and raw data for qualitative and quantitative analysis are available at https://github.com/daisyinb612/CO-OPERA.

cs.HC

Argus: Federated Non-convex Bilevel Learning over 6G Space-Air-Ground Integrated Network

The space-air-ground integrated network (SAGIN) has recently emerged as a core element in the 6G networks. However, traditional centralized and synchronous optimization algorithms are unsuitable for SAGIN due to infrastructureless and time-varying environments. This paper aims to develop a novel Asynchronous algorithm a.k.a. Argus for tackling non-convex and non-smooth decentralized federated bilevel learning over SAGIN. The proposed algorithm allows networked agents (e.g. autonomous aerial vehicles) to tackle bilevel learning problems in time-varying networks asynchronously, thereby averting stragglers from impeding the overall training speed. We provide a theoretical analysis of the iteration complexity, communication complexity, and computational complexity of Argus. Its effectiveness is further demonstrated through numerical experiments.

cs.LG

Hierarchical Equivariant Policy via Frame Transfer

Recent advances in hierarchical policy learning highlight the advantages of decomposing systems into high-level and low-level agents, enabling efficient long-horizon reasoning and precise fine-grained control. However, the interface between these hierarchy levels remains underexplored, and existing hierarchical methods often ignore domain symmetry, resulting in the need for extensive demonstrations to achieve robust performance. To address these issues, we propose Hierarchical Equivariant Policy (HEP), a novel hierarchical policy framework. We propose a frame transfer interface for hierarchical policy learning, which uses the high-level agent's output as a coordinate frame for the low-level agent, providing a strong inductive bias while retaining flexibility. Additionally, we integrate domain symmetries into both levels and theoretically demonstrate the system's overall equivariance. HEP achieves state-of-the-art performance in complex robotic manipulation tasks, demonstrating significant improvements in both simulation and real-world settings.

cs.RO

Equivariant Diffusion Policy

Recent work has shown diffusion models are an effective approach to learning the multimodal distributions arising from demonstration data in behavior cloning. However, a drawback of this approach is the need to learn a denoising function, which is significantly more complex than learning an explicit policy. In this work, we propose Equivariant Diffusion Policy, a novel diffusion policy learning method that leverages domain symmetries to obtain better sample efficiency and generalization in the denoising function. We theoretically analyze the $\mathrm{SO}(2)$ symmetry of full 6-DoF control and characterize when a diffusion model is $\mathrm{SO}(2)$-equivariant. We furthermore evaluate the method empirically on a set of 12 simulation tasks in MimicGen, and show that it obtains a success rate that is, on average, 21.9% higher than the baseline Diffusion Policy. We also evaluate the method on a real-world system to show that effective policies can be learned with relatively few training samples, whereas the baseline Diffusion Policy cannot.

cs.RO

StableGarment: Garment-Centric Generation via Stable Diffusion

In this paper, we introduce StableGarment, a unified framework to tackle garment-centric(GC) generation tasks, including GC text-to-image, controllable GC text-to-image, stylized GC text-to-image, and robust virtual try-on. The main challenge lies in retaining the intricate textures of the garment while maintaining the flexibility of pre-trained Stable Diffusion. Our solution involves the development of a garment encoder, a trainable copy of the denoising UNet equipped with additive self-attention (ASA) layers. These ASA layers are specifically devised to transfer detailed garment textures, also facilitating the integration of stylized base models for the creation of stylized images. Furthermore, the incorporation of a dedicated try-on ControlNet enables StableGarment to execute virtual try-on tasks with precision. We also build a novel data engine that produces high-quality synthesized data to preserve the model's ability to follow prompts. Extensive experiments demonstrate that our approach delivers state-of-the-art (SOTA) results among existing virtual try-on methods and exhibits high flexibility with broad potential applications in various garment-centric image generation.

cs.CV

Shear-induced droplet mobility within porous surfaces

Droplet mobility under shear flows is important in a wide range of engineering applications, e.g., fog collection, and self-cleaning surfaces. For structured surfaces to achieve superhydrophobicity, the removal of stains adhered within the microscale surface features strongly determines the functional performance and durability. In this study, we numerically investigate the shear-induced mobility of the droplet trapped within porous surfaces. Through simulations covering a wide range of flow conditions and porous geometries, three droplet mobility modes are identified, i.e., the stick-slip, crossover, and slugging modes. To quantitatively characterise the droplet dynamics, we propose a droplet-scale capillary number that considers the driving force and capillary resistance. By comparing against the simulation results, the proposed dimensionless number presents a strong correlation with the leftover volume. The dominating mechanisms revealed in this study provide a basis for further research on enhancing surface cleaning and optimising design of anti-fouling surfaces.

physics.flu-dyn

Matter flow method for alleviating checkerboard oscillations in triangular mesh SGH Lagrangian simulation

When the SGH Lagrangian based on triangle mesh is used to simulate compressible hydrodynamics, because of the stiffness of triangular mesh, the problem of physical quantity cell-to-cell spatial oscillation (also called "checkerboard oscillation") is easy to occur. A matter flow method is proposed to alleviate the oscillation of physical quantities caused by triangular stiffness. The basic idea of this method is to attribute the stiffness of triangle to the fact that the edges of triangle mesh can not do bending motion, and to compensate the effect of triangle edge bending motion by means of matter flow. Three effects are considered in our matter flow method: (1) transport of the mass, momentum and energy carried by the moving matter; (2) the work done on the element, since the flow of matter changes the specific volume of the grid element; (3) the effect of matter flow on the strain rate in the element. Numerical experiments show that the proposed matter flow method can effectively alleviate the spatial oscillation of physical quantities.

physics.comp-ph

New Open Cluster Candidates Discovered in the XSTPS-GAC Survey

The Xuyi Schmidt Telescope Photometric Survey of the Galactic Anti-center (XSTPS-GAC) is a photometric sky survey that covers nearly 6 000 deg^2 towards Galactic anti-center in g r i bands. Half of its survey field locates on the Galactic Anti-center disk, which makes XSTPS-GAC highly suitable for searching new open clusters in the GAC region. In this paper, we report new open cluster candidates discovered in this survey, as well as properties of these open cluster candidates, such as age, distance and reddening, derived by isochrone fitting in the color-magnitude diagram (CMD). These open cluster candidates are stellar density peaks detected in the star density maps by applying the method from Koposov et al. (2008). Each candidate is inspected on its true color image composed from XSTPS-GAC three band images. Then its CMD is checked, in order to identify whether the central region stars have a clear isochrone-like trend differing from the background stars. The parameters derived from isochrone fitting for these candidates are mainly based on three band photometry of XSTPS-GAC. Meanwhile, when these new candidates are able to be seen clearly on 2MASS, their parameters are also derived based on 2MASS (J-H, J) CMD. Finally, there are 320 known open clusters rediscovered and 24 new open cluster candidates discovered in this work. Further more, the parameters of these new candidates, as well as another 11 known recovered open clusters, are properly determined for the first time.

astro-ph.SR