SearcharxivSearch

arXiv subjects

Ziyi Zhang

Publications and source records attributed to Ziyi Zhang.

At least 19 recordsLinked to original sources

Learning Fractional-Order Dynamics from a Single Trajectory

Many real-world processes exhibit long-range dependence, where the current state depends on a slowly decaying trace of past states rather than on the most recent state alone. This paper studies system identification for discrete-time fractional-order linear time-invariant systems from a single observed trajectory of length $t$, a setting that captures such non-Markovian dynamics through the Grünwald--Letnikov difference operator. Unlike Markovian systems, fractional-order systems couple estimation across the entire history, making both statistical analysis and practical identification more challenging. We propose \emph{Fractional-Order Ordinary-Least-Squares Grid-Search (FO-GS)}, a simple two-stage estimator that exploits the diagonal structure of the fractional-difference operator to decouple the identification problem row-wise. Under the stability assumption, we establish high-probability, non-asymptotic error bounds for estimating both the fractional order and the system matrix in the heterogeneous setting, with both estimation errors scaling as \(\mathcal{O}(t^{-1/2})\). Through experiments, we show that \emph{FO-GS} outperforms existing baselines in recovering both the fractional order and the underlying system dynamics.

cs.LG

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories, and simulated experience. Across 28 embodied understanding benchmarks, our 8B model achieves an average score of 72.5, setting a new open-source state of the art and performing on par with leading proprietary models such as GPT-6-Astra and Gemini 3.6 Flash. It achieves the best open-source results on 14 benchmarks while retaining general multimodal capabilities. Beyond these understanding evaluations, qualitative examples show the model's ability to produce end-effector trajectories and predict future scenes through spatially aligned RGB, depth, and robot-mask outputs.

cs.CV

Boundary Cases of the $J$-Equation: Divisorial Rigidity and a Global $C^0$ Estimate

We study two boundary cases of the stability condition for the $J$-equation. First, under the $J$-semistable condition, we show that the destabilizing prime divisors form an exceptional family in the sense of Boucksom, with a uniform numerical gap away from them. The related modified nef estimate can remove the $J$-big assumption in Liu's work~\cite{Liu2026Boundary}. Second, under the smooth boundary cone condition, we obtain a uniform global $C^0$ estimate for solutions of the approximating twisted $J$-equations. As a consequence, we construct a bounded-potential Bedford--Taylor solution of the $J$-equation, which is smooth outside the destabilizing prime divisors.

math.DG

Classification of Novikov-Poisson Algebras and Their Applications

In this paper, we give a complete classification, up to isomorphism, of 3-dimensional complex Novikov--Poisson algebras. As an application of the classification, we further prove that every 3-dimensional complex transposed Poisson algebra can be obtained from a Novikov--Poisson algebra except the Lie algebra $\mathfrak{sl}_2(\mathbb C)$. Consequently, Sartayev's conjecture holds in dimension 3, that is, every 3-dimensional complex transposed Poisson algebra is special when regarded as a GD algebra.

math.RA

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a lightweight post-hoc corrector. Prior automated feature engineering models the data-generating process; corrective features instead model the model-failure process. We present CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model (LLM) that proposes named feature combinations, binary flags, and short executable code. A single validation-grounded gate accepts or rejects every candidate regardless of its origin, and a validation-selected corrector applies the accepted features or leaves the forecast unchanged. This source-agnostic pipeline also allows prior feature-engineering systems to be evaluated under identical conditions, making CRAFTER an instrument for attributing forecast improvements to the feature source alone. Across six public datasets and six frozen backbones, CRAFTER surpasses every dedicated feature-engineering system at every feature budget, roughly doubling the improvement achieved by the corrector alone and reducing the error of the weakest backbones by up to 27%. These gains are robust across different LLM backends and persist even when applied on top of fine-tuned backbones.

cs.LG

A Numerical Criterion for the 2-Hessian Equation on Compact Kähler Manifolds

We show that a Nakai--Moishezon-type criterion associated with the complex $2$-Hessian equation produces a Gauduchon class. In complex dimension three, this numerical criterion is equivalent to the existence of a smooth $2$-admissible representative and hence to the solvability of the $2$-Hessian equation. As consequences of these results, we prove the corresponding conjectures of Murakami for the complex Hessian equation and of Székelyhidi for the Hessian quotient equation in dimension three. We also establish a boundary version of the above results.

math.DG

A CubeSat Electronics System for Dual-Satellite Coordinated Soft X-Ray Polarimetry

Addressing the unique requirements for a wide field of view and rapid response in soft X-ray polarization measurements of transient sources such as gamma-ray bursts, this paper proposes the design of a high-reliability CubeSat payload electronics system for dual-satellite cooperative observation. The system employs a Gas Microchannel Pixel Detector (GMPD) and a Topmetal-L sensor as its core components, forming a highly integrated, low-noise payload hardware platform. It achieves a field of view of 90$^\circ$ $\times$ 90$^\circ$, a sensitive area of 3.69 cm$^{2}$, and a power consumption of less than 6 W, operating within an energy range of 2--10 keV. The system incorporates autonomous high-voltage (HV) ramp-up/ramp-down control and a dual-protection mechanism based on count rate and discharge events, enabling in-orbit responses to risks such as the South Atlantic Anomaly and solar particle events. It also supports single-event upset detection and recovery, as well as in-orbit firmware upgrades. The communication interface adopts a redundant primary/backup Controller Area Network (CAN) bus design, with measured channel switching times of less than 100 ms, meeting the demands for real-time command interaction in dual-satellite coordination. By utilizing a large-array pixel sensor with region-of-interest readout and integrating an in-orbit track compression algorithm, the system significantly reduces data storage and downlink transmission resource burdens. Ground tests demonstrate an equivalent noise charge of 22.35 e$^{-}$, HV monitoring linearity better than $\pm$0.5\%, and an output range extending to -5 kV. Thermal vacuum cycling tests show no performance degradation after five cycles between -5 $^\circ$C and 40 $^\circ$C. This work demonstrates the system's capability for autonomous observation, intelligent coordination, and reliable operation in complex space environments.

physics.ins-det

TUGS: Physics-based Compact Representation of Underwater Scenes by Tensorized Gaussian

Underwater 3D scene reconstruction is crucial for multimedia applications in adverse environments, such as underwater robotic perception and navigation. However, the complexity of interactions between light propagation, water medium, and object surfaces poses significant difficulties for existing methods in accurately simulating their interplay. Additionally, expensive training and rendering costs limit their practical application. Therefore, we propose Tensorized Underwater Gaussian Splatting (TUGS), a compact underwater 3D representation based on physical modeling of complex underwater light fields. TUGS includes a physics-based underwater Adaptive Medium Estimation (AME) module, enabling accurate simulation of both light attenuation and backscatter effects in underwater environments, and introduces Tensorized Densification Strategies (TDS) to efficiently refine the tensorized representation during optimization. TUGS is able to render high-quality underwater images with faster rendering speeds and less memory usage. Extensive experiments on real-world underwater datasets have demonstrated that TUGS can efficiently achieve superior reconstruction quality using a limited number of parameters. The code is available at https://liamlian0727.github.io/TUGS

cs.CV

Gelfand--Dorfman Algebras: Nilpotency, Solvability, Construction and Classification

In this paper, we characterize the nilpotency and solvability of Gelfand--Dorfman (GD) algebras. In contrast with Poisson algebras and transposed Poisson algebras, we give examples show the nilpotency and solvability of a GD algebra are not determined by the nilpotency and solvability of its underlying algebras. To obtain more examples of special and non-special GD algebras, we give several construction methods and determine whether the resulting algebras are special. Futhermore, we study GD algebra structures on simple Lie algebras. We provide examples demonstrating that GD algebra structures on simple Lie algebras are not necessarily trivial, distinguishing them from Poisson and transposed Poisson algebras. Finally, we provide a complete algebraic classification of low-dimensional complex GD algebras, and determine their nilpotency, solvability and speciality.

math.RA

SAHG: Sector-Anisotropic Hyperbolic Graph Model for Social Bot Detection

LLM-driven social bots can generate fluent, human-like text, reducing the discriminative advantage of content-based detection alone. However, coordinated campaigns still leave relational patterns -- interactions, behavioral similarity, shared neighborhoods, community positions, and coordinated activity -- that graph-based methods can exploit. Existing graph detectors face two challenges when exploiting such evidence. First, Euclidean GNNs distort hierarchical and scale-free social graphs; while hyperbolic geometry addresses this volume-growth mismatch, fixed-curvature models still assign uniform geometric resolution to structural directions with different densities and separation needs. Second, relational evidence is not always reliable: sophisticated bots forge heterophilic connections with genuine users, causing neighborhood aggregation to mix bot and human signals and dilute account-level evidence. We propose SAHG (Sector-Anisotropic Hyperbolic Graph), addressing both challenges. SAHG learns a direction-dependent curvature field $γ(u)$ that adapts geometric resolution across structural directions, and uses sector prototypes to convert angular concentration and alignment into classifier-readable features. To prevent contaminated aggregation from overwhelming account-level evidence, SAHG encodes per-account features and graph-neighborhood representations in two independent SAH channels, fusing them only at the classifier. Experiments on Fox8-23, BotSim-24, and MGTAB show that SAHG achieves the highest accuracy and F1 on all three benchmarks, outperforming feature-based, graph-based, LLM-based, and isotropic hyperbolic baselines. Ablation and geometric analyses confirm the effectiveness of the anisotropic geometry and dual-channel design.

cs.SI

$A$-Generalized Hessian pre-Lie algebras and $A$-Generalized Yang--Baxter Equations

Inspired by the problem of constructing ($ω$-)pre-Lie algebra structures on the dual space of a pre-Lie algebra, we introduce the \(A\)-generalized Yang--Baxter equation as a generalization of the Yang--Baxter equation of pre-Lie algebras. We study its symmetric solutions through \(A\)-generalized Hessian pre-Lie algebras and split these solutions into two types. We further consider factorizable solutions of this equation and establish a one-to-one correspondence between them and generalized quadratic Rota--Baxter pre-Lie algebras of nonzero weight. By studying the structure of these algebras, we find all factorizable solutions. Finally, we study the structure of \(A\)-generalized Hessian pre-Lie algebras. In particular, we obtain a structural description via central and double extensions and classify low-dimensional non-trivial \(A\)-generalized Hessian pre-Lie algebras.

math-ph

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a cognitively inspired benchmark that evaluates Omni-MLLMs across three stages, perception, understanding, and reasoning, through cross-modal tasks requiring joint audio-visual interpretation. This design enables fine-grained diagnosis of model capabilities and failure modes. To further assess robustness beyond familiar domains, we propose AVI-Bench-PriSe, an extension that probes models' primitive audio-visual sensation using unfamiliar, low-semantic stimuli, testing generalization beyond common training distributions. Extensive experiments on both open-source and closed-source models reveal substantial limitations in current Omni-MLLMs. Based on these findings, we present a four-level AVI taxonomy. Overall, AVI-Bench provides a principled evaluation framework to guide the development of more robust and generalizable AVI. Project website: https://fudancvl.github.io/AVI-Bench/

cs.CV

Exploring Accuracy Law for Deep Time Series Forecasters: An Empirical Study

Deep time series forecasting has emerged as a rapidly growing field in recent years. Despite the exponential growth of community interests, progress on standard benchmarks is often limited to marginal improvements. A common consensus of the community is that time series forecasting inherently faces a non-zero error lower bound due to its partially observable and uncertain nature. However, a fundamental question arises: how to estimate the performance upper bound of deep time series forecasters? We delve into univariate time series forecasting, a prevalent forecasting paradigm spanning traditional statistical models to advanced time series foundation models. Going beyond classical series-wise predictability metrics, we realize that the forecasting performance is highly related to window-wise properties due to the sequence-to-sequence forecasting paradigm of deep time series models and introduce a quantitative measurement of window-wise pattern complexity. Through rigorous statistical analyses over more than 4700 newly trained deep forecasting models, we discover a consistent empirical relationship between the minimum attainable forecasting error of deep models and the complexity of window-wise series patterns, which is termed the accuracy law. We further demonstrate that this empirical finding successfully guides us to identify saturated tasks from widely used benchmarks and derive an effective training strategy for time series foundation models, offering valuable insights for future research.

cs.LG

NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science

The automation of scientific research workflows has emerged as a transformative frontier in artificial intelligence, yet existing autonomous research agents remain largely domain-agnostic, lacking the specialized reasoning, method selection, and data acquisition capabilities required for rigorous spatial data science. This paper introduces NORA (Night Owl Research Agent), a harness-engineered, multi-agent autonomous research system purpose-built for GIScience and spatial data science. NORA orchestrates the complete research lifecycle through a skills-first architecture comprising 21 domain-specialized workflow skills, 9 specialist sub-agents, and custom Model Context Protocol (MCP) servers. Central to the system's design are two novel domain-specialized skills: a spatial analysis skill unit that encodes decision frameworks for exploratory spatial data analysis, spatial regression, and diagnostics; and a spatial data download skill that supports reproducible acquisition from authoritative geospatial data sources. We formalize the concept of harness engineering for scientific research agents, demonstrating how lifecycle hooks, safety gates, generator-evaluator separation, human-in-the-loop, and state persistence ensure reliable and reproducible autonomous research. We evaluate NORA through case studies by 6 domain specialists and 3 LLM reviewers across seven dimensions (novelty, quality, rigor, etc). Results demonstrate that domain-specialized harness engineering substantially improves the efficiency and quality of research output compared to general-purpose agent configurations.

cs.AI

Reducing Detail Hallucinations in Long-Context Regulatory Understanding via Targeted Preference Optimization

Large language models (LLMs) frequently produce \emph{detail hallucinations} when processing long regulatory documents, including subtle errors in threshold values, units, scopes, obligation levels, and conditions that preserve surface plausibility while corrupting safety-critical parameters. We formalize this phenomenon through a fine-grained \emph{Detail Error Taxonomy} of five error types and introduce \textbf{DetailBench}, a benchmark built from 172 real regulatory documents and 150 synthetic documents spanning three jurisdictions, with human-annotated detail-level ground truth comprising 13,000 preference pairs. We propose \textbf{DetailDPO}, a targeted preference optimization framework that constructs contrastive pairs differing in exactly one detail dimension, concentrating DPO gradient signal on detail-bearing~tokens. We provide theoretical analysis showing why \emph{minimal detail perturbation} pairs yield gradient concentration under mild assumptions. Experiments on the Qwen2.5 family (7B, 14B, 72B) and Llama-3.1-8B across three context-length tiers (8K--64K tokens) show that DetailDPO reduces the Detail Error Rate by 42--61\% relative to baselines, with consistent gains across all five error types and cross-domain transfer to financial and medical documents.

cs.SI

An unfitted finite element method for PDE-constrained shape optimization via shape gradient flow

In this paper, we propose an unfitted finite element method to solve PDE-constrained shape optimization problems via shape gradient flow. The shape gradient flow system consists of the state equation, the adjoint equation, the velocity equation, as well as the flow map that generates the evolution of the boundary driven by the velocity field, which can be viewed as a limit system of the classical shape gradient descent algorithm. In \cite{GongLiRao} the authors proposed an evolving finite element method to solve the shape gradient flow system. Instead, in this paper, we propose an unfitted finite element method in which the evolution of the boundary is realized by cubic splines and the equations are solved by cut finite element methods with ghost penalization. Under reasonable assumptions, we are able to prove some optimal convergence rates that are further validated by numerical experiments.

math.NA

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consistency. Reinforcement learning (RL) finetuning offers a potential solution, yet existing approaches designed for single-image diffusion do not readily extend to the few-step T2MV setting, as they neglect cross-view coordination and suffer from weak learning signals in few-step regimes. To address this, we propose MVC-ZigAL, a tailored RL finetuning framework for few-step T2MV diffusion models. Specifically, its core insights are: (1) a new MDP formulation that jointly models all generated views and assesses their collective quality via a joint-view reward; (2) a novel advantage learning strategy that exploits the performance gains of a self-refinement sampling scheme over standard sampling, yielding stronger learning signals for effective RL finetuning; and (3) a unified RL framework that extends advantage learning with a Lagrangian dual formulation for multiview-constrained optimization, balancing single-view and joint-view objectives through adaptive primal-dual updates under a self-paced threshold curriculum that harmonizes exploration and constraint enforcement. Collectively, these designs enable robust and balanced RL finetuning for few-step T2MV diffusion models, yielding substantial gains in both per-view fidelity and cross-view consistency. Code is available at https://github.com/ZiyiZhang27/MVC-ZigAL.

cs.LG

Telogenesis: Goal Is All U Need

Goal-conditioned systems assume goals are provided externally. We ask whether attentional priorities can emerge endogenously from an agent's internal cognitive state. We propose a priority function that generates observation targets from three epistemic gaps: ignorance (posterior variance), surprise (prediction error), and staleness (temporal decay of confidence in unobserved variables). We validate this in two systems: a minimal attention-allocation environment (2,000 runs) and a modular, partially observable world (500 runs). Ablation shows each component is necessary. A key finding is metric-dependent reversal: under global prediction error, coverage-based rotation wins; under change detection latency, priority-guided allocation wins, with advantage growing monotonically with dimensionality (d = -0.95 at N=48, p < 10^-6). Detection latency follows a power law in attention budget, with a steeper exponent for priority-guided allocation (0.55 vs. 0.40). When the decay rate is made learnable per variable, the system spontaneously recovers environmental volatility structure without supervision (t = 22.5, p < 10^-6). We demonstrate that epistemic gaps alone, without external reward, suffice to generate adaptive priorities that outperform fixed strategies and recover latent environmental structure.

cs.AI