SearcharxivSearch

arXiv subjects

Keru Zhou

Publications and source records attributed to Keru Zhou.

8 recordsLinked to original sources

EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy

A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation and information exchange. A Main Agent retains task-level state within an active context window, while role-specific Sub Agents process perception, execution monitoring, verification, and memory consolidation in isolated contexts and return structured, task-relevant evidence. Role-specific contexts control information load by exposing only decision-relevant evidence to the Main Agent, while the functional Skill interface composes heterogeneous backends as Operational, Imagination, and Evaluation Skills. Criterion-grounded verification, textual failure diagnosis, and Branch Stack recovery provide localized correction, with token-aware external memory preserving task-relevant state. Together, their closed-loop interaction realizes the system-level policy captured by the name EMERGE-Policy. Without additional fine-tuning, we achieved outstanding performance on several public benchmark that have had a wide-reaching impact, and conducted a series of real robot experiments. These system-level results suggest that through the division of different functional sub-tasks among multiple agents and their concurrent collaboration, as well as the technical paradigm where the model is regarded as a skill and called within the framework, EMERGE-Policy can extend the robust robot policies beyond isolated runs.

cs.RO

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation

Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segmentation masks encode only the former. We present KAM-WM, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning. KAM-WM queries a Flow Matching image-to-video backbone once and interprets its single-step latent velocity as a Kinematic Affordance Map (KAM), which provides task-conditioned interaction regions and coarse motion structure. A lightweight Perceiver compresses KAM into tokens that condition a diffusion policy together with RGB observations and proprioception. Across LIBERO and RoboTwin2.0, KAM-WM reaches 90.6% average success on LIBERO and achieves 65.7% and 22.4% success rates in the Easy and Hard settings on RoboTwin2.0, respectively. Controlled comparisons against a zero-order mask prior suggest that part of the gains comes from directional information beyond spatial localization alone. These results indicate that, in the evaluated settings, a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.

cs.RO

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observations, without explicitly modeling future dynamics. This limits their ability to generalize under out-of-distribution perturbations. To address this issue, we propose ELAN4D, an embodiment-centric, 4D-aware training framework that enhances VLA policies with future robot keypoint tracks as predictive spatio-temporal supervision. Using only forward kinematics from proprioceptive states, we derive 3D displacement tracks of robot keypoints, such as joints and the end-effector, with negligible preprocess cost. These tracks provide metric and compact supervision without requiring external trackers or reconstruction. A plug-and-play auxiliary branch with a lightweight track decoder injects this 4D signal into the action expert while preserving the pretrained vision-language backbone through gradient isolation. The track decoder is discarded during inference, leaving the base policy interface unchanged. Extensive experiments on LIBERO, LIBERO-Plus, RoboTwin2.0 and real-world manipulation tasks demonstrate that ELAN4D consistently improves over strong VLA baselines, achieving the best overall performance and substantial gains under out-of-distribution perturbations, including camera, background, and layout shifts. These results highlight the effectiveness of embodiment-centric 4D supervision for building more robust and generalizable manipulation policies.

cs.RO

Sparsity-Exploiting Channel Estimation For Unsourced Random Access With Fluid Antenna

This work explores the channel estimation (CE) problem in uplink transmission for unsourced random access (URA) with a fluid antenna receiver. The additional spatial diversity in a fluid antenna system (FAS) addresses the needs of URA design in multiple-input and multiple-output (MIMO) systems. We present two CE strategies based on the activation of different FAS ports, namely alternate ports and partial ports CE. Both strategies facilitate the estimation of channel coefficients and angles of arrival (AoAs). Additionally, we discuss how to refine channel estimation by leveraging the sparsity of finite scatterers. Specifically, the proposed partial ports CE strategy is implemented using a regularized estimator, and we optimize the estimator's parameter to achieve the desired AoA precision and refinement. Extensive numerical results demonstrate the feasibility of the proposed strategies, and a comparison with a conventional receiver using half-wavelength antennas highlights the promising future of integrating URA and FAS.

cs.IT

Two Families of Constant Term Identities

In 1985, Bressoud and Goulden derived the formula for the constant term in $\prod_{(i,j)\in T} \frac{x_j}{x_i}\\\prod_{0\le i<j \le n}(\frac{x_i}{x_j})_{a_i}(\frac{qx_j}{x_i})_{a_j-1}$, where $T \subseteq \{(i,j)\mid 0\le i<j \le n\}$. This result implies the Andrews' $q$-Dyson identity. In 2006, Gessel and Xin proved the $q$-Dyson identity by considering both sides of the equality as polynomials in $q^{a_0}$. We use this approach to determine the coefficients of $x_0/x_1$ and $x_0/x_2$ in Laurent polynomials studied by Bressoud and Goulden.

math.CO

$q$-Fractional Integral Operators With Two Parameters

We use the Poisson kernel of the continuous $q$-Hermite polynomials to introduces families of integral operators, which are semigroups of linear operators. We describe the eigenvalues and eigenfunctions of one family of operators. The action of the semigroups of operators on the Askey--Wilson polynomials is shown to only change the parameters but preserves the degrees, hence we produce transmutation relation for the Askey--Wilson polynomials. The transmutation relations are then used to derive bilinear generating functions involving the Askey--Wilson polynomials.

math.CA

Orthogonal Polynomials of Askey-Wilson Type

We study two families of orthogonal polynomials. The first is a finite family related to the Askey-Wilson polynomials but the orthogonality is on the real line. A limiting case of this family is an infinite system of orthogonal polynomials whose moment problem is indeterminate. We provide several orthogonality measures for the infinite family and derive their Plancherel-Rotach asymptotics.

math.CA

q-Fractional Askey-Wilson Integrals and Related Semigroups of Operators

We introduce three one-parameter semigroups of operators and determine their spectra. Two of them are fractional integrals associated with the Askey-Wilson operator. We also study these families as families of positive linear approximation operators. Applications include connection relations and bilinear formulas for the Askey-Wilson polynomials. We also introduce a q-Gauss-Weierstrass transform and prove a representation and inversion theorem for it.

math.CA