SearcharxivSearch

arXiv subjects

Zhenyuan Zhang

Publications and source records attributed to Zhenyuan Zhang.

At least 19 recordsLinked to original sources

MemCoRe: Recovering Evidence from Progressively Compressed Factual Knowledge for Agent Memory

Memory systems enable LLM agents to consolidate and retrieve relevant evidence from the factual knowledge accumulated through growing interaction histories for downstream reasoning. Existing approaches have explored diverse strategies for organizing and compressing these histories. However, balancing compression with retrieval effectiveness remains challenging: retaining too much content can cause relevant evidence to be obscured by redundant entries, while discarding too aggressively may remove content that later proves relevant. This amounts to a tradeoff between compressing redundancy and preserving enough structure to retrieve target evidence, as formalized by the information bottleneck. To this end, we propose MemCoRe, which organizes memory as a compression hierarchy where each level compresses redundancy further while retaining the structure needed for retrieval at that level. In this hierarchy, evidence is progressively compressed from detailed records through extracted keywords to topic groups. This enables retrieval to locate target evidence by searching across levels of the hierarchy. Comprehensive experiments demonstrate that MemCoRe outperforms existing state-of-the-art baselines.

cs.AI

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.

cs.AI

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

cs.AI

The exact dimensional threshold for Spearman rank-correlation compatibility

We show that, for a given dimension, the set of Spearman's rank correlation matrices and that of linear correlation matrices coincide if and only if the dimension is no larger than nine. For this, we construct an extreme rank-four counterexample in dimension ten and prove its incompatibility using moment identities and Cauchy-Schwarz. Appending unit directions produces counterexamples in every higher dimension. This, together with existing results, completes the dimensional classification and settles a long-standing open question in quantitative risk management.

math.ST

Precise cover times for branching random walks on Hamming graphs: (iterated) logarithmic corrections

We prove tight asymptotics of the cover time $τ_{\mathrm{cov}}(d)$ of a continuous-time branching random walk on the Hamming graph $\{0,1,\dots,b-1\}^d$, as $d\to\infty$. We focus on the slow-branching regime, where particles move at rate one and branch at rate $λ\in(0,1)$. For $b>2$, we show that $τ_{\mathrm{cov}}(d)=x_\star d+λ^{-1}\log d+O_{\mathbb P}(1)$. For $b=2$, we show that $τ_{\mathrm{cov}}(d)=x_\star d+χ^{-1}\log\log d+O_{\mathbb P}(1)$. Here, $x_\star$ and $χ$ are explicit positive constants depending only on $b$ and $λ$. Our results sharpen previously known linear-order estimates. The dichotomy reflects the geometry of the last uncovered region: for $b>2$, there are exponentially many antipodes, whereas the binary hypercube has a unique antipode and its neighbors govern the final coverage. Our proofs combine classic spine change of measure techniques and many-to-few estimates with a multiscale decomposition of the genealogy and a weighted martingale analysis of the early population.

math.PR

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

Autonomous data collection governs the volume and quality of real-world trajectories for manipulation policy learning. Existing pipelines reduce human effort via self-resetting, VLM verification, or language-guided correction, yet episode-scoped fixes must be reissued whenever the same failure recurs, so oversight cost grows with session length rather than with the number of distinct problems. We present Zero2Skill, a human-robot symbiotic agentic system in which corrections are retained and reused across rounds. The collection loop collects, verifies, and resets autonomously, pausing for a remote operator only when a phase exhausts an explicit retry budget. An LLM parser maps each natural-language utterance to a structured adjustment stored in Corrective Memory, so addressed failure modes typically need not be corrected again under the same conditions. On a real-robot desktop-clearing testbed, Zero2Skill matches teleoperation episode success while reducing human working time to 16%. Language corrections improve verifier-human agreement in all four evaluated settings and raise average single-attempt success from 12.5% to 47.5% (arm-selection: 20.0% to 50.0%). Policies fine-tuned on Zero2Skill data match teleoperation-trained policy success at a fraction of collection human cost.

cs.RO

A Control Theory of Predictability in Latent World Models

Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current practice adopts the prediction error, the single- or multi-step rollout loss on held-out data, as the training and model-selection objective, on the assumption that a lower prediction error yields better control. We show that this assumption is unreliable for a structural reason: a planner does not query the model on the training distribution but on the states that its candidate actions reach, which generally leave the data manifold, so an error averaged over the data cannot by itself govern control. We therefore reframe the objective as the discrepancy between the predicted and the true plan-cost at the plan the planner commits to, and prove that the planner's suboptimality is bounded by twice this discrepancy, whereas the data-averaged prediction error neither bounds nor tracks it. Under a linear-control premise the discrepancy separates into two terms. The first is a small on-manifold residual, on which the predicted and true dynamics agree and which a spectral tax prices through the non-normality of the latent transition operator. The second is an off-manifold divergence, on which an action carries the state off the manifold and the two dynamics diverge; this divergence is the binding term and is bounded by no data-averaged error. Synthetic operators confirm the pricing formulas, and latent model-predictive control experiments confirm the decoupling: across seeds, the single-step validation error is essentially uncorrelated with control success, whereas a fidelity score on the planner-reachable measure tracks it.

cs.LG

Decorated stable $p$-adic self-similar processes with stationary increments

We construct new classes of examples of self-similar processes with stationary increments indexed by $\mathbb Q_p$ via stable integrals. Classical constructions arise from the real counterpart and from discounted branching random walks. We discuss a new decoration technique that significantly enlarges these classes. The decoration technique makes use of the special symmetry of $\mathbb{Q}_p$ to obtain self-similarity and stationarity of increments, and it does not have an analogue on the real line. We also show that these enlarged classes of decorated processes are pairwise incomparable under inclusion.

math.PR

Consensus on Dynamic Stochastic Block Models: Fast Convergence and Phase Transitions

We introduce two models of consensus following a majority rule on time-evolving stochastic block models (SBM), in which the network evolution is Markovian or non-Markovian. Under the majority rule, in each round, each agent simultaneously updates their opinion according to the majority of their neighbors. Our network has a community structure and randomly evolves with time. In contrast to the classic setting, the dynamics is not purely deterministic, and reflects the structure of SBM by resampling the connections at each step, making agents with the same opinion more likely to connect than those with different opinions. In the Markovian model, connections between agents are resampled at each step according to the SBM law and each agent updates their opinion via the majority rule. We prove a power-of-one type result, i.e., any initial bias leads to a non-trivial advantage of winning in the end, uniformly in the size of the network. In the non-Markovian model, a connection between two agents is resampled according to the SBM law only when at least one of them changes opinion and is otherwise kept the same. We identify the phase-transition threshold, up to the second-order leading term, between halting and fast convergence to consensus. We also give sufficient initial-lead conditions for consensus to occur within one, two, or three rounds.

math.PR

Sample Path Properties of the Fractional Wiener--Weierstrass Bridge II

Fractional Wiener--Weierstrass bridges are a class of Gaussian processes obtained by replacing trigonometric functions in the construction of classical Weierstrass functions by fractional Brownian bridges. A number of their sample path properties were derived in Schied--Zhang (2024,2026). The analysis in these papers left several open questions, most of which are addressed here. Specifically, we prove that, in the regime in which the Weierstrass mechanism dominates the underlying fractional Brownian bridge, the limiting $b$-adic variation coefficient has an absolutely continuous distribution and is therefore genuinely random. At the critical point between the two roughness regimes, we establish the power-variation formula and the critical $Φ$-variation limit conjectured in Schied--Zhang (2024). Finally, we derive the Hausdorff dimension for the graphs of the sample paths by proving a conjecture from Schied--Zhang (2026) for the missing high-Hurst case.

math.PR

Diamond transports in quadratic-form and distorted optimal transport

The diamond transport is generated by the uniform law on a diamond-shaped copula support. Since a classical optimal transport (OT) objective is affine in the coupling, this transport cannot be the unique minimizer in the classical setting. We study a broader family of transports, called diamond-type transports, in non-classical settings such as quadratic-form optimal transport (QOT) and distorted optimal transport (DOT), which are generally nonconvex. Our main results are within the QOT framework: for symmetric one-dimensional marginals, the diamond transport is an optimizer for a large class of QOT problems whose costs depend on within-coordinate distances. Examples include product costs under positive-definiteness and convexity conditions and, in particular, mixed rectangular costs. For rectangular costs, we show that the diamond transport is the unique minimizer except for boundary cases. In the DOT framework, diamond-type transports are minimizers for a natural class of cost, and the diamond transport is the unique minimizer in specialized examples. We also identify the intersection between DOT and QOT, which corresponds precisely to quadratic distortion functions.

math.OC

Quadratic-form Optimal Transport

We introduce the framework of quadratic-form optimal transport (QOT), whose transport cost has the form $\iint c\,\mathrm{d}π\otimes\mathrm{d}π$ for some coupling $π$ between two marginals. Interesting examples of quadratic-form transport cost and their optimization include inequality measurement, the variance of a bivariate function, covariance, Kendall's tau, the Gromov--Wasserstein distance, quadratic assignment problems, and quadratic regularization of classic optimal transport. QOT leads to substantially different mathematical structures compared to classic transport problems and many technical challenges. We illustrate the fundamental properties of QOT and provide several cases where explicit solutions are obtained. For a wide class of cost functions, including the rectangular cost functions, the QOT problem is solved by a new coupling called the diamond transport, whose copula is supported on a diamond in the unit square.

math.PR

Almost periodicity as a path property for $p$-adic self-similar processes with stationary increments

Shen and Zhang (2021) showed that almost periodicity naturally arises in the spectral representation of discrete-time $p$-adic self-similar processes with stationary increments. In this paper, we study several notions of almost periodicity as sample path properties of Banach space-valued $p$-adic sssi processes. We prove that Bohr almost periodicity is equivalent, as a path event, to continuity with respect to the $p$-adic topology. We also show that the corresponding equivalence fails for Weyl and Besicovitch almost periodicity. Finally, we extend the Bohr almost-periodic result to finite-dimensional random fields.

math.PR

ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation

Current AI agent frameworks have made remarkable progress in automating individual tasks, yet all existing systems serve a single user. Human productivity rests on the social and organizational relationships through which people coordinate, negotiate, and delegate. When agents move beyond performing tasks for one person to representing that person in collaboration with others, the infrastructure for cross-user agent collaboration is entirely absent, let alone the governance mechanisms needed to secure it. We argue that the next frontier for AI agents lies not in stronger individual capability, but in the digitization of human collaborative relationships. To this end, we propose a human-symbiotic agent paradigm. Each user owns a permanently bound agent system that collaborates on the owner's behalf, forming a network whose nodes are humans rather than agents. This paradigm rests on three governance primitives. A layered identity architecture separates a Manager Agent from multiple context-specific Identity Agents; the Manager Agent holds global knowledge but is architecturally isolated from external communication. Scoped authorization enforces per-identity access control and escalates boundary violations to the owner. Action-level accountability logs every operation against its owner's identity and authorization, ensuring full auditability. We instantiate this paradigm in ClawNet, an identity-governed agent collaboration framework that enforces identity binding and authorization verification through a central orchestrator, enabling multiple users to collaborate securely through their respective agents.

cs.AI

Stabilizing the Splits through Minimax Decision Trees

By revisiting the end-cut preference (ECP) phenomenon associated with a single CART (Breiman et al. (1984)), we introduce MinimaxSplit decision trees, a robust alternative to CART that selects splits by minimizing the worst-case child risk rather than the average risk. For regression, we minimize the maximum within-child squared error; for classification, we minimize the maximum child entropy, yielding a C4.5-compatible criterion. We also study a cyclic variant that deterministically cycles coordinates, leading to our main method of cyclic MinimaxSplit decision trees. We prove oracle inequalities that cover both regression and classification, under mild marginal non-atomicity conditions. The bounds control the tree's global excess risk by local worst-case impurities and yield fast convergence rates compared to CART. We extend the analysis to a random-dimension forest variant that subsamples coordinates per node. Empirically, (cyclic) MinimaxSplit trees and their forests improve over baselines on structured heterogeneous data such as EEG amplitude regression over fixed time horizons and image denoising, framed as non-parametric regression on spatial coordinates.

math.ST

Tightness Analysis of First Passage Times of $d$-Dimensional Branching Random Walk

Given a discrete-time non-lattice supercritical branching random walk in $\mathbb{R}^d$, we investigate its first passage time to a shifted unit ball of a distance $x$ from the origin, conditioned upon survival. We provide precise asymptotics up to $O(1)$ (tightness) for the first passage time as a function of $x$ as $x\to\infty$, thus resolving a conjecture in Blanchet--Cai--Mohanty--Zhang (2024). Our proof builds on the previous analysis of Blanchet--Cai--Mohanty--Zhang (2024) and employs a careful multi-scale analysis on the genealogy of particles within a distance of $\asymp \log x$ near extrema of a one-dimensional branching random walk, where the cluster structure plays a crucial role.

math.PR

Viral Quasispecies Evolution as a Branching Random Walk on the Hypercube

We study a continuous-time nearest-neighbor branching random walk on the $d$-dimensional $b$-ary hypercube $\{0,1,\dots,b-1\}^d$ as a model for viral quasispecies evolution under mutation and replication. Motivated by mutagenic antiviral treatments and evolutionary-safety questions, we analyze the first passage time to a fixed target genotype at Hamming distance $m$, corresponding to the first appearance of a prescribed collection of mutations. We derive sharp asymptotics for these first passage times, uniformly for $m\le d/L$ as $d\to\infty$ (where $L>0$ is a large constant), and identify a phase transition in first-passage scaling at $ρ=e$, where $ρ$ denotes the effective growth parameter. In the slow-branching regime $ρ\in(1,e)$ relevant to mutagenic treatment scenarios, the first passage time is asymptotically affine in the genome length $d$ and the target distance $m$. In particular, when replication is fixed and mutation exceeds branching, increasing the mutation rate can delay the first appearance of a prescribed genotype by order $d$, providing a quantitative perspective on evolutionary safety.

math.PR

Boundedness of discounted branching random walks via generic chaining

Consider a discrete-time supercritical discounted branching random walk, in which increments at depth $k$ are independent and identically distributed with the same law as $m^{-kH}Y$, where $Y$ has a fixed law, $H>0$, and $m>1$ is the expected number of offspring at depth one. We provide a clean characterization of the boundedness of the discounted branching random walk: under mild conditions on the offspring distribution, the process is almost surely bounded if and only if $\mathbb{E}[|Y|^{1/H}]<\infty$. This extends results of Athreya (1985) and Aïdékon--Hu--Shi (2024), and provides a partial answer to Open Problem 31 of Aldous--Bandyopadhyay (2005).

math.PR