SearcharxivSearch

arXiv subjects

Zezhou Zhang

Publications and source records attributed to Zezhou Zhang.

10 recordsLinked to original sources

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-person video: before an explicit customer request, an agent must use limited human-object interaction evidence to intervene or remain silent. Physical grounding here means converting observations into task-relevant retail state, not modeling low-level dynamics. We introduce the Proactive Intent World Model (PIWM): See constructs the perceptual basis, Foresee models counterfactual consequences, and Act selects an action. Performance is poor when the agent must extract information from raw video and decide directly, but improves substantially with structured inputs extracted and annotated from a professional retail perspective. AIDA-stage constraints and BDI-state ablations further support role- and goal-directed selection and organization of decision-relevant cues. Counterfactual prediction performs well in standalone evaluation, yet planning methods that query these forecasts at inference time degrade sharply: locally useful consequence prediction does not reliably improve action selection. This gap may reflect incomplete process understanding, uncertainty in fine-grained single-step outcomes, and insufficient joint modeling of scenes and temporal evolution. Hold remains the hardest action in structured-state evaluation, exposing a related challenge in temporal awareness. PIWM advances static intent recognition toward intent world modeling by organizing observations under task knowledge, anticipating candidate interventions, and treating intervention and non-intervention jointly. Future work will introduce long-horizon interaction trajectories and temporal consequence supervision to improve sustained reasoning and intervention timing.

cs.CL

Learning Graph-Indexed Trajectory Patterns for Stochastic On-Time Arrival Routing

Correlated link travel times create decision-relevant patterns in partial route histories. In stochastic on-time arrival (SOTA) routing, each route prefix forms a variable-length, graph-indexed sequence in which traversed-edge identities, realized travel times, and route order jointly indicate the reliability of downstream actions. We present GPG-HT, a history-conditioned Transformer policy that learns a trajectory representation from this structured sequence together with the current node, destination, and remaining budget. Edge-time cross-attention and sequence encoding capture dependencies within the observed history, while decoder cross-attention maps the resulting trajectory memory and decision context to an online distribution over feasible outgoing edges. A history-conditioned generalized policy-gradient objective trains the representation from terminal on-time outcomes. Experiments on the Sioux Falls and Anaheim road-network topologies with simulated correlated link times show that GPG-HT achieves higher mean on-time arrival probabilities than representative optimization and reinforcement-learning baselines. Paired common-pool evaluation confirms statistically significant gains in all six network-budget settings, reaching 2.82-3.27 percentage points on Sioux Falls and 0.36-1.04 percentage points on Anaheim. Correlated, independent, shuffled-history, no-history, and architecture controls further demonstrate that GPG-HT learns decision-relevant structure from graph-indexed route prefixes.

cs.LG

Superconformal algebras over arbitrary rings of coefficients

We define analogs of superconformal algebras over an arbitrary commutative associative superalgebra. In the finitely generated case the universal central extensions of these algebras are finitely presented. In particular, we show that all known superconformal algebras are finitely presented.

math.RA

PVI: Plug-in Visual Injection for Vision-Language-Action Models

VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for semantic abstraction and typically conditioned on static visual observations, tends to attenuate fine-grained geometric cues and often lacks explicit temporal evidence for the action expert. Prior work mitigates this by injecting auxiliary visual features, but existing approaches either focus on static spatial representations or require substantial architectural modifications to accommodate temporal inputs, leaving temporal information underexplored. We propose Plug-in Visual Injection (PVI), a lightweight, encoder-agnostic module that attaches to a pretrained action expert and injects auxiliary visual representations via zero-initialized residual pathways, preserving pretrained behavior with only single-stage fine-tuning. Using PVI, we obtain consistent gains over the base policy and a range of competitive alternative injection strategies, and our controlled study shows that temporal video features (V-JEPA2) outperform strong static image features (DINOv2), with the largest gains on multi-phase tasks requiring state tracking and coordination. Real-robot experiments on long-horizon bimanual cloth folding further demonstrate the practicality of PVI beyond simulation.

cs.CV

Cyclic homology of Jordan superalgebras and related Lie superalgebras

We study the relationship between cyclic homology of Jordan superalgebras and second cohomologies of their Tits-Kantor-Koecher Lie superalgebras. In particular, we focus on Jordan superalgebras that are Kantor doubles of bracket algebras. The obtained results are applied to computation of second cohomologies and universal central extensions of Hamiltonian and contact type Lie superalgebras over arbitrary rings of coefficients.

math.RA

Diffusion Probabilistic Model Based Accurate and High-Degree-of-Freedom Metasurface Inverse Design

Conventional meta-atom designs rely heavily on researchers' prior knowledge and trial-and-error searches using full-wave simulations, resulting in time-consuming and inefficient processes. Inverse design methods based on optimization algorithms, such as evolutionary algorithms, and topological optimizations, have been introduced to design metamaterials. However, none of these algorithms are general enough to fulfill multi-objective tasks. Recently, deep learning methods represented by Generative Adversarial Networks (GANs) have been applied to inverse design of metamaterials, which can directly generate high-degree-of-freedom meta-atoms based on S-parameter requirements. However, the adversarial training process of GANs makes the network unstable and results in high modeling costs. This paper proposes a novel metamaterial inverse design method based on the diffusion probability theory. By learning the Markov process that transforms the original structure into a Gaussian distribution, the proposed method can gradually remove the noise starting from the Gaussian distribution and generate new high-degree-of-freedom meta-atoms that meet S-parameter conditions, which avoids the model instability introduced by the adversarial training process of GANs and ensures more accurate and high-quality generation results. Experiments have proven that our method is superior to representative methods of GANs in terms of model convergence speed, generation accuracy, and quality.

cs.LG

Finite presentability of universal central extensions of ${\mathfrak{sl}_n}$

In this paper we discuss finite presentability of the universal central extensions of Lie algebras ${\mathfrak{sl}_n(R)}$, where $n\geq 3$ and $R$ is a unital associative $k$-algebra. We show that a universal central extension is finitely presented if and only if the algebra $R$ is finitely presented.

math.RA

Finite presentability of universal central extensions of ${\mathfrak{sl}_n}$, II

In this note we connect finite presentability of a Jordan algebra to finite presentability of its Tits-Kantor-Koecher algebra. Through this we complete our discussion of finite presentability of universal central extensions of ${\mathfrak{sl}_n(A)}$, $A$ a $k$-algebra, initiated in our previous paper, and answer a question raised by Shestakov-Zelmanov in the positive.

math.RA

Property (T) for Kac-Moody groups over rings

Let R be a finitely generated commutative ring with 1, let A be an indecomposable 2-spherical generalized Cartan matrix of size at least 2 and M=M(A) the largest absolute value of a non-diagonal entry of A. We prove that there exists an integer n=n(A) such that the Kac-Moody group G_A(R) has property (T) whenever R has no proper ideals of index less than n and all positive integers less than or equal to M are invertible in R.

math.GR