SearcharxivSearch

arXiv subjects

Kaiyi Xiong

Publications and source records attributed to Kaiyi Xiong.

2 recordsLinked to original sources

Population-Scale Advancing Interface Modeling Reveals How Bacterial Swarms Encode Future Spatial Architecture

Motile bacteria shape microbial function by occupying space, yet how collective motion becomes population-scale architecture remains poorly resolved. Bacterial swarming is not merely surface motion, but a process by which motile populations commit to future macroscopic form. Here, in Enterobacter sp. SM3, a gut-associated swarmer linked to mucosal repair, we treat the advancing colony--environment interface as a morphodynamic state through which local motility becomes spatial order. We built SwarmEvo across thermal, hydration, and substrate-mechanical conditions and developed Morpher to resolve and propagate interface states. Counterintuitively, within the permissive assay range, condition labels only weakly separated future trajectories, whereas colony-specific interface geometry constrained later expansion, indicating that swarm fate is written into the interface rather than prescribed by condition identity. Boundary fidelity was decisive: a 0.67 percentage-point segmentation gap expanded into a 2.4--3.1 IoU-point forecasting loss. Preserving front displacement, protrusion continuity, and branch memory, Morpher predicted late-stage expansion with 95.42% mIoU, 10.61 px HD$_{95}$, and 3.93 px ASSD. These results identify the advancing interface as a state-bearing layer through which motility and environmental constraint are converted into future spatial form, enabling disease-relevant microbial organization to be read before endpoint architecture emerges.

cond-mat.soft

GazeVLA: Learning Human Intention for Robotic Manipulation

Embodied foundation models have achieved significant breakthroughs in robotic manipulation, yet they still depend heavily on large-scale robot demonstrations. Although recent works have explored leveraging human data to alleviate this dependency, effectively extracting transferable knowledge remains a significant challenge due to the inherent embodiment gap between human and robot. We argue that the intention underlying human actions can serve as a powerful intermediate representation for bridging this gap. In this paper, we introduce a novel framework that explicitly learns and transfers human intention to facilitate robotic manipulation. Specifically, we model intention through gaze, as it naturally precedes physical actions and serves as an observable proxy for human intent. Our model is first pretrained on a large-scale egocentric human dataset to capture human intention and its synergy with action, followed by finetuning on a small set of robot and human data. During inference, the model adopts a Chain-of-Thought reasoning paradigm, sequentially predicting intention before executing the action. Extensive evaluations in simulation and real-world settings, across long-horizon and fine-grained tasks, and under few-shot and robustness benchmarks, show that our method consistently outperforms strong baselines, generalizes better, and achieves state-of-the-art performance. Project page: https://gazevla.github.io .

cs.RO