SearcharxivSearch

arXiv subjects

Yongchang Li

Publications and source records attributed to Yongchang Li.

8 recordsLinked to original sources

Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration

Foundation models have become central to unifying perception and planning in robotics, yet real-world deployment exposes a mismatch between their monolithic assumption that a single model can handle all cognitive functions and the distributed, dynamic nature of practical service workflows. Vision-language models offer strong semantic understanding but lack embodiment-aware action capabilities while relying on hand-crafted skills. Vision-Language-Action policies enable reactive manipulation but remain brittle across embodiments, weak in geometric grounding, and devoid of proactive collaboration mechanisms. These limitations indicate that scaling a single model alone cannot deliver reliable autonomy for service robots operating in human-populated settings. To address this gap, we present InteractGen, an LLM-powered multi-agent framework that decomposes robot intelligence into specialized agents for continuous perception, dependency-aware planning, decision and verification, failure reflection, and dynamic human delegation, treating foundation models as regulated components within a closed-loop collective. Deployed on a heterogeneous robot team and evaluated in a three-month open-use study, InteractGen improves task success, adaptability, and human-robot collaboration, providing evidence that multi-agent orchestration offers a more feasible path toward socially grounded service autonomy than further scaling standalone models.

cs.RO

CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human

In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxiliary generative models, by integrating VLM-based reflective reasoning with diffusion-based action generation under a mixture-of-experts design. Through a two-stage training recipe of action grounding and reflection tuning, it supports explicit self-reflection and proactively solicits human guidance when confronted with uncertainty or repeated failure. It cuts normalized Time by ~2x and Dream counts by ~4x vs. generative agents, achieving higher success rates, improved interpretability, and balanced low latency compared with existing methods. This work takes a pioneering step toward shifting VLAs from opaque controllers to genuinely assistive agents capable of reasoning, acting, and collaborating with humans.

cs.RO

Charge density wave with suppressed long-range structural modulation in canted antiferromagnetic kagome FeGe

Kagome lattice can host abundant exotic quantum states such as superconductivity and charge density wave (CDW). Recently, successive orders of A-type antiferromagnetism (AFM), CDW and canted AFM have been manifested upon cooling in kagome FeGe. However, the mechanism of CDW and interaction with magnetism remains unclear. Here we investigate the evolution of CDW with temperature across the canted AFM by single-crystal x-ray diffraction, scanning tunneling microscope (STM) and resonant elastic x-ray scattering (REXS). Interestingly, CDW-induced superlattice reflections become weak after the canted AFM, although long-range CDW order is still detectable by STM and REXS. We uncover a novel long-range CDW order with suppressed structural modulation, likely due to the competition for the underlying crystal structure between CDW and canted AFM. Additionally, occupational modulations of Ge1 in the kagome plane and displacive modulations of all atoms were extracted. The results confirm Ge dimerization along the c axis and suggest a dynamic transformation between different CDW domains.

cond-mat.str-el

AssistantX: An LLM-Powered Proactive Assistant in Collaborative Human-Populated Environment

Current service robots suffer from limited natural language communication abilities, heavy reliance on predefined commands, ongoing human intervention, and, most notably, a lack of proactive collaboration awareness in human-populated environments. This results in narrow applicability and low utility. In this paper, we introduce AssistantX, an LLM-powered proactive assistant designed for autonomous operation in realworld scenarios with high accuracy. AssistantX employs a multi-agent framework consisting of 4 specialized LLM agents, each dedicated to perception, planning, decision-making, and reflective review, facilitating advanced inference capabilities and comprehensive collaboration awareness, much like a human assistant by your side. We built a dataset of 210 real-world tasks to validate AssistantX, which includes instruction content and status information on whether relevant personnel are available. Extensive experiments were conducted in both text-based simulations and a real office environment over the course of a month and a half. Our experiments demonstrate the effectiveness of the proposed framework, showing that AssistantX can reactively respond to user instructions, actively adjust strategies to adapt to contingencies, and proactively seek assistance from humans to ensure successful task completion. More details and videos can be found at https://assistantx-agent.github.io/AssistantX/.

cs.RO

Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble

As real-world knowledge is constantly evolving, ensuring the timeliness and accuracy of a model's knowledge is crucial. This has made knowledge editing in large language models increasingly important. However, existing knowledge editing methods face several challenges, including parameter localization coupling, imprecise localization, and a lack of dynamic interaction across layers. In this paper, we propose a novel knowledge editing method called Knowledge Neuronal Ensemble (KNE). A knowledge neuronal ensemble represents a group of neurons encoding specific knowledge, thus mitigating the issue of frequent parameter modification caused by coupling in parameter localization. The KNE method enhances the precision and accuracy of parameter localization by computing gradient attribution scores for each parameter at each layer. During the editing process, only the gradients and losses associated with the knowledge neuronal ensemble are computed, with error backpropagation performed accordingly, ensuring dynamic interaction and collaborative updates among parameters. Experimental results on three widely used knowledge editing datasets show that the KNE method significantly improves the accuracy of knowledge editing and achieves, or even exceeds, the performance of the best baseline methods in portability and locality metrics.

cs.CL

Limits of dispersoid size and number density in oxide dispersion strengthened alloys fabricated with powder bed fusion-laser beam

Previous work on additively-manufactured oxide dispersion strengthened alloys focused on experimental approaches, resulting in larger dispersoid sizes and lower number densities than can be achieved with conventional powder metallurgy. To improve the as-fabricated microstructure, this work integrates experiments with a thermodynamic and kinetic modeling framework to probe the limits of the dispersoid sizes and number densities that can be achieved with powder bed fusion-laser beam. Bulk samples of a Ni-20Cr $+$ 1 wt.% Y$_2$O$_3$ alloy are fabricated using a range of laser power and scanning velocity combinations. Scanning transmission electron microscopy characterization is performed to quantify the dispersoid size distributions across the processing space. The smallest mean dispersoid diameter (29 nm) is observed at 300 W and 1200 mm/s, with a number density of 1.0$\times$10$^{20}$ m$^{-3}$. The largest mean diameter (72 nm) is observed at 200 W and 200 mm/s, with a number density of 1.5$\times$10$^{19}$ m$^{-3}$. Scanning electron microscopy suggests that a considerable fraction of the oxide added to the feedstock is lost during processing, due to oxide agglomeration and the ejection of oxide-rich spatter from the melt pool. After accounting for these losses, the model predictions for the dispersoid diameter and number density align with the experimental trends. The results suggest that the mechanism that limits the final number density is collision coarsening of dispersoids in the melt pool. The modeling framework is leveraged to propose processing strategies to limit dispersoid size and increase number density.

cond-mat.mtrl-sci

Multipiezo effect in altermagnetic V2SeTeO monolayer

Inspired by recent theoretical proposal on the interesting piezomagnetism and C-paired valley polarization in V2Se2O monolayer, we predict a stable antiferromagnetic Janus monolayer V2SeTeO with altermagnetic configuration using density functional theory calculations. It exhibits a novel multi-piezo effect combining piezoelectric, piezovalley and piezomagnetism. Most interestingly, the valley polarization and the net magnetization under strain in V2SeTeO exceed these in V2Se2O, along with the additional large piezoelectric coefficient of e31 (0.322*10-10 C m-1). The multi-piezo effect makes antiferromagnetic Janus monolayer V2SeTeO a tantalizing material for potential applications in nanoelectronics, optoelectronics, spintronics and valleytronics.

cond-mat.mtrl-sci

Two-dimensional charge density wave TaX$_2$ (X=S, Se, Te) from first principles

Transition metal dichalcogenides are rich in their structural phases, e.g. 1T-TaS2 and 1T-TaSe2 form charge density wave (CDW) under low temperature with interesting and exotic properties. Here, we present a systematic study of different structures in two-dimensional TaX2 (X=S, Se, Te) using density functional theory calculations with consideration of van der Waals interaction. All the normal phases present metal characteristics with various ground state and magnetic properties. The lattice reconstruction of CDW drastically affects the electronic and structural characteristics of 1T-TaS2 and 1T-TaSe2, leading to a transition from metal to insulator and an emergence of magnetic moment within periodic atomic clusters called the Star of David. The evaluated Heisenberg couplings indicate the weak ferromagnetic coupling between the clusters in monolayer. Furthermore, in bilayer commensurate CDW cases, we find intriguing phenomenon of the varying magnetic properties with different stacking orders. The magnetic moment in each layer disappears when two layers are coupled, but may sustain in certain stackings of interlayer antiferromagnetic configurations.

cond-mat.mtrl-sci