SearcharxivSearch

arXiv subjects

Zhipeng Liu

Publications and source records attributed to Zhipeng Liu.

At least 19 recordsLinked to original sources

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Captions serve as a primary supervision signal for both multimodal understanding and text-to-image generation. However, previous evaluations treat the caption quality as a single scalar objective, which conflates two distinct properties: (1) how much visual information a caption covers and (2) how reliably the image supports its stated claims. To this end, we design a decoupled caption evaluation benchmark, CAPEval (Coverage And Precision Evaluation), with human-written ground-truth captions and human-verified atomic checklist items. Specifically, CAPEval decomposes caption quality into Coverage and Precision. The former quantifies how thoroughly a caption covers ground-truth factual content, while the latter reflects the factual correctness rate of all claims expressed in the caption. We select 10 captioners and further conduct controlled downstream end-to-end experiments with them from four model families, where the caption source is the only variable. Empirically, we find a consistent task-dependent dissociation: Coverage serves as the stronger correlate for understanding performance, whereas Precision acts as the dominant predictor for generation performance. This decoupled evaluation paradigm not only delivers a more fine-grained diagnosis of caption quality, but also offers actionable guidance for selecting and optimizing captioners tailored to different downstream tasks.

cs.CV

CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection

Vision-language models (VLMs) enable text-guided object detection but degrade severely under cross-view scenarios where ground and aerial viewpoints differ in altitude, scale, and spatial layout. These geometric changes introduce systematic complexity variations between viewpoints, e.g., ground view images contain dense and highly occluded structures, while aerial images are sparse and globally organized. Fixed VLM fusion mechanisms cannot handle this discrepancy. We propose CrossVL, a framework combining Complexity-Aware Pathway Aggregation (CPA) and Paired Curriculum Learning (PCL) for enhanced cross-view detection for VLM. CPA estimates scene complexity from multimodal statistics and routes visual features through multiple pathways to obtain view-specific representations. PCL leverages semantic consistency of synchronized ground-aerial pairs to provide stable early supervision and then gradually shifts toward randomized sampling. On MAVREC, CrossVL improves Florence-2's aerial mAP from 58.66% to 61.03% and reduces the ground-aerial performance gap from 8.63pp to 6.65pp, while also achieving a 3.3x reduction in variance across random seeds. CPA provides stable complexity-aware feature aggregation, and PCL enhances optimization dynamics. Together, they demonstrate that coordinated architectural and training adaptations are crucial for robust cross-view VLM detection.

cs.CV

CombinationTS: A Modular Framework for Understanding Time-Series Forecasting Models

Recent progress in time-series forecasting has led to rapidly increasing architectural complexity, yet many reported State-of-the-Art gains are statistically fragile or misattributed. We argue that progress requires a shift from model selection to modular attribution, identifying which components truly drive performance. We propose CombinationTS, a self-contained probabilistic evaluation framework that decomposes forecasting models into orthogonal modules--Input Transformation, Embedding, Encoder, Decoder, and Output Transformation--and evaluates them under a shared evaluation condition space. By quantifying each component via marginalized performance ($\mu$) and stability ($\sigma$), CombinationTS enables robust attribution beyond fragile point estimates. Through large-scale paired evaluation, we uncover the Identity Paradox: once the data view (Embedding) is well-designed, a parameter-free Identity Encoder often matches or outperforms complex backbones. We further show that explicit structural priors introduced via Input Transformations yield a more favorable performance-stability trade-off than increasing Encoder complexity, establishing a principled baseline for architectural necessity.

cs.LG

A determinant identity for the sum of contour integral matrices

We derive an identity for the determinant of the sum of two $n\times n$ matrices, $U$ and $M$, whose entries are defined via contour integrals. Specifically, we consider $U(i,j)=\frac{1}{2\pi\mathrm{i}}\oint_{\mathrm{C}} \frac{\prod_{\ell=1}^{i-1} (z-\beta_\ell)}{\prod_{\ell=1}^{j} (z-\beta_\ell)} p_i(z)f_j(z)\mathrm{d} z$ and $M(i,j)= \frac{1}{2\pi \mathrm{i}}\int_{\Gamma} q_i(z)g_j(z) \mathrm{d} z$. Under suitable assumptions on the functions $p,q,f,g$, we show that $\det(U+M)$ can be expressed as a Fredholm determinant $\det(\mathrm{I} +K)$, where $K$ is an integral kernel acting on the contour $\Gamma$. The kernel $K$ depends on a function $H$ that solves a system of integral equations. When $f_i$ and $g_i$ are specialized to certain rational functions depending on two sets of parameters $(\alpha_\ell)_{\ell\in \mathbb{Z}}$ and $(\beta_\ell)_{\ell\in \mathbb{Z}}$, $H$ becomes the characteristic function associated with inhomogeneous directed last passage percolation (DLPP) and inhomogeneous totally asymmetric simple exclusion process (TASEP) models. Furthermore, we obtain an explicit random walk hitting expectation representation of this characteristic function. Our work generalizes a recent identity by Baik, Liao, and Liu (2026), which plays an important role in finding the multipoint distribution formula of the periodic KPZ fixed point. Finally, we demonstrate three applications of our general formulas in integrable probability: a new Fredholm determinant formula for the distribution of the path-to-point last passage time in the inhomogeneous DLPP, an indirect proof of a new path-to-line joint distribution formula in the homogeneous DLPP, and a novel proof of the TASEP path-integral formula previously obtained by Matetski, Quastel, and Remenik (2021).

math.CA

MorphSNN: Adaptive Graph Diffusion and Structural Plasticity for Spiking Neural Networks

Spiking Neural Networks (SNNs) currently face a critical bottleneck: while individual neurons exhibit dynamic biological properties, their macro-scopic architectures remain confined within conventional connectivity patterns that are static and hierarchical. This discrepancy between neuron-level dynamics and network-level fixed connectivity eliminates critical brain-like lateral interactions, limiting adaptability in changing environments. To address this, we propose MorphSNN, a backbone framework inspired by biological non-synaptic diffusion and structural plasticity. Specifically, we introduce a Graph Diffusion (GD)mechanism to facilitate efficient undirected signal propagation, complementing the feedforward hierarchy. Furthermore, it incorporates a Spatio-Temporal Structural Plasticity (STSP) mechanism, endowing the network with the capability for instance-specific, dynamic topological reorganization, thereby overcoming the limitations of fixed topologies. Experiments demonstrate that MorphSNN achieves state-of-the-art accuracy on static and neuromorphic datasets; for instance, it reaches 83.35% accuracy on N-Caltech101 with only 5 timesteps. More importantly, its self-evolving topology functions as an intrinsic distribution fingerprint, enabling superior Out-of- Distribution (OOD) detection without auxiliary training. The code is available at anonymous.4open.science/r/MorphSNN-B0BC.

cs.NE

Periodic KPZ fixed point with general initial conditions

We consider the relaxation-time-scale limit of the periodic totally asymmetric simple exclusion process (PTASEP) with general initial conditions. For every sequence of initial conditions approximating a periodic upper semicontinuous function, we compute the limiting space-time multipoint distributions of the rescaled particle locations and height functions. The resulting finite-dimensional distributions are explicit and form a consistent family, thereby defining a spatially periodic space-time random field. We call this field the periodic KPZ fixed point with the corresponding initial condition. This extends earlier results for PTASEP with special initial conditions and defines the periodic analogue of the KPZ fixed point on the line. The main technical novelty is a pair of new probabilistic representations for the energy function and the characteristic function, the two functions through which the initial condition enters the finite-time PTASEP multipoint distribution formula. Both representations are expressed in terms of a geometric random walk and two stopping times, namely the first hitting time of the initial profile and the first such hitting time at or after one full period, with the latter capturing the periodic geometry.

math.PR

Data-Centric Benchmark for Label Noise Estimation and Ranking in Remote Sensing Binary Building Segmentation

High-quality pixel-level annotations are essential for the semantic segmentation of remote sensing imagery. However, such labels are expensive to obtain and often affected by noise due to the labor-intensive and time-consuming nature of pixel-wise annotation, which makes it challenging for human annotators to label every pixel accurately. Annotation errors can significantly degrade the performance and robustness of modern segmentation models, motivating the need for reliable mechanisms to identify and quantify noisy training samples. This paper introduces a novel data-centric benchmark, together with a new, publicly available binary building segmentation dataset. This specific task serves as a representative and practically relevant testbed that enables controlled experimentation with different annotation perturbations. Furthermore, we also introduce two techniques for identifying, quantifying, and ranking training samples according to their level of label noise in remote sensing semantic segmentation. Such proposed methods leverage complementary strategies based on model uncertainty, prediction consistency, and representation analysis, and consistently outperform established baselines across a range of experimental settings. The outcomes of this work are publicly available at https://github.com/keillernogueira/label_noise_segmentation.

cs.CV

VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics

While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our work, we introduce Video Connecting, an innovative task that aims to generate smooth intermediate video content between given start and end clips. However, the absence of standardized evaluation benchmarks has hindered the development of this task. To bridge this gap, we proposed VC-Bench, a novel benchmark specifically designed for video connecting. It includes 1,579 high-quality videos collected from public platforms, covering 15 main categories and 72 subcategories to ensure diversity and structure. VC-Bench focuses on three core aspects: Video Quality Score VQS, Start-End Consistency Score SECS, and Transition Smoothness Score TSS. Together, they form a comprehensive framework that moves beyond conventional quality-only metrics. We evaluated multiple state-of-the-art video generation models on VC-Bench. Experimental results reveal significant limitations in maintaining start-end consistency and transition smoothness, leading to lower overall coherence and fluidity. We expect that VC-Bench will serve as a pioneering benchmark to inspire and guide future research in video connecting. The evaluation metrics and dataset are publicly available at: https://anonymous.4open.science/r/VC-Bench-1B67/.

cs.CV

Beyond Visual Realism: Toward Reliable Financial Time Series Generation

Generative models for financial time series often create data that look realistic and even reproduce stylized facts such as fat tails or volatility clustering. However, these apparent successes break down under trading backtests: models like GANs or WGAN-GP frequently collapse, yielding extreme and unrealistic results that make the synthetic data unusable in practice. We identify the root cause in the neglect of financial asymmetry and rare tail events, which strongly affect market risk but are often overlooked by objectives focusing on distribution matching. To address this, we introduce the Stylized Facts Alignment GAN (SFAG), which converts key stylized facts into differentiable structural constraints and jointly optimizes them with adversarial loss. This multi-constraint design ensures that generated series remain aligned with market dynamics not only in plots but also in backtesting. Experiments on the Shanghai Composite Index (2004--2024) show that while baseline GANs produce unstable and implausible trading outcomes, SFAG generates synthetic data that preserve stylized facts and support robust momentum strategy performance. Our results highlight that structure-preserving objectives are essential to bridge the gap between superficial realism and practical usability in financial generative modeling.

q-fin.ST

We Need a More Robust Classifier: Dual Causal Learning Empowers Domain-Incremental Time Series Classification

The World Wide Web thrives on intelligent services that rely on accurate time series classification, which has recently witnessed significant progress driven by advances in deep learning. However, existing studies face challenges in domain incremental learning. In this paper, we propose a lightweight and robust dual-causal disentanglement framework (DualCD) to enhance the robustness of models under domain incremental scenarios, which can be seamlessly integrated into time series classification models. Specifically, DualCD first introduces a temporal feature disentanglement module to capture class-causal features and spurious features. The causal features can offer sufficient predictive power to support the classifier in domain incremental learning settings. To accurately capture these causal features, we further design a dual-causal intervention mechanism to eliminate the influence of both intra-class and inter-class confounding features. This mechanism constructs variant samples by combining the current class's causal features with intra-class spurious features and with causal features from other classes. The causal intervention loss encourages the model to accurately predict the labels of these variant samples based solely on the causal features. Extensive experiments on multiple datasets and models demonstrate that DualCD effectively improves performance in domain incremental scenarios. We summarize our rich experiments into a comprehensive benchmark to facilitate research in domain incremental time series classification.

cs.LG

Transformation Journey of Zr-based MOFs: Study on Mechanics and Hydrogen Storage under Doping Regulation

This study delves into the transformation journey of Zr-based Metal-Organic Frameworks (MOFs), focusing on enhancing their mechanical properties and hydrogen storage capacities through doping regulation. MOFs, a versatile class of crystalline porous materials, have garnered significant attention due to their unique properties and broad potential applications in gas storage, separation, catalysis, and sensing. Among them, Zr-based MOFs stand out for their exceptional stability and high surface area. This research systematically investigates six key Zr-based MOFs (UIO-66, UIO-67, UIO-68, MOF-801, MOF-802, and MOF-841) using multiscale computational methods, including molecular dynamics (MD) simulations, grand canonical Monte Carlo (GCMC) simulations, and density functional theory (DFT). The study explores the impact of metal ion substitution (Fe, Co, Ni, Cu, Zn) on the mechanical and hydrogen storage properties of these MOFs. Our findings reveal that metal ion substitution significantly influences the mechanical stability and hydrogen adsorption capacity of Zr-based MOFs, providing valuable insights for the design and optimization of high-performance MOF materials.

cond-mat.mtrl-sci

CogniSNN: Enabling Neuron-Expandability, Pathway-Reusability, and Dynamic-Configurability with Random Graph Architectures in Spiking Neural Networks

Spiking neural networks (SNNs), regarded as the third generation of artificial neural networks, are expected to bridge the gap between artificial intelligence and computational neuroscience. However, most mainstream SNN research directly adopts the rigid, chain-like hierarchical architecture of traditional artificial neural networks (ANNs), ignoring key structural characteristics of the brain. Biological neurons are stochastically interconnected, forming complex neural pathways that exhibit Neuron-Expandability, Pathway-Reusability, and Dynamic-Configurability. In this paper, we introduce a new SNN paradigm, named Cognition-aware SNN (CogniSNN), by incorporating Random Graph Architecture (RGA). Furthermore, we address the issues of network degradation and dimensional mismatch in deep pathways by introducing an improved pure spiking residual mechanism alongside an adaptive pooling strategy. Then, we design a Key Pathway-based Learning without Forgetting (KP-LwF) approach, which selectively reuses critical neural pathways while retaining historical knowledge, enabling efficient multi-task transfer. Finally, we propose a Dynamic Growth Learning (DGL) algorithm that allows neurons and synapses to grow dynamically along the internal temporal dimension. Extensive experiments demonstrate that CogniSNN achieves performance comparable to, or even surpassing, current state-of-the-art SNNs on neuromorphic datasets and Tiny-ImageNet. The Pathway-Reusability enhances the network's continuous learning capability across different scenarios, while the dynamic growth algorithm improves robustness against interference and mitigates the fixed-timestep constraints during neuromorphic chip deployment. This work demonstrates the potential of SNNs with random graph structures in advancing brain-inspired intelligence and lays the foundation for their practical application on neuromorphic hardware.

cs.NE

Gated Fusion Enhanced Multi-Scale Hierarchical Graph Convolutional Network for Stock Movement Prediction

Accurately predicting stock market movements remains a formidable challenge due to the inherent volatility and complex interdependencies among stocks. Although multi-scale Graph Neural Networks (GNNs) hold potential for modeling these relationships, they frequently neglect two key points: the subtle intra-attribute patterns within each stock affecting inter-stock correlation, and the biased attention to coarse- and fine-grained features during multi-scale sampling. To overcome these challenges, we introduce MS-HGFN (Multi-Scale Hierarchical Graph Fusion Network). The model features a hierarchical GNN module that forms dynamic graphs by learning patterns from intra-attributes and features from inter-attributes over different time scales, thus comprehensively capturing spatio-temporal dependencies. Additionally, a top-down gating approach facilitates the integration of multi-scale spatio-temporal features, preserving critical coarse- and fine-grained features without too much interference. Experiments utilizing real-world datasets from U.S. and Chinese stock markets demonstrate that MS-HGFN outperforms both traditional and advanced models, yielding up to a 1.4% improvement in prediction accuracy and enhanced stability in return simulations. The code is available at https://anonymous.4open.science/r/MS-HGFN.

cs.LG

TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting

Although Transformers excel in natural language processing, their extension to time series forecasting remains challenging due to insufficient consideration of the differences between textual and temporal modalities. In this paper, we develop a novel Transformer architecture designed for time series data, aiming to maximize its representational capacity. We identify two key but often overlooked characteristics of time series: (1) unidirectional influence from the past to the future, and (2) the phenomenon of decaying influence over time. These characteristics are introduced to enhance the attention mechanism of Transformers. We propose TimeFormer, whose core innovation is a self-attention mechanism with two modulation terms (MoSA), designed to capture these temporal priors of time series under the constraints of the Hawkes process and causal masking. Additionally, TimeFormer introduces a framework based on multi-scale and subsequence analysis to capture semantic dependencies at different temporal scales, enriching the temporal dependencies. Extensive experiments conducted on multiple real-world datasets show that TimeFormer significantly outperforms state-of-the-art methods, achieving up to a 7.45% reduction in MSE compared to the best baseline and setting new benchmarks on 94.04\% of evaluation metrics. Moreover, we demonstrate that the MoSA mechanism can be broadly applied to enhance the performance of other Transformer-based models.

cs.LG

TimeExpert: Boosting Long Time Series Forecasting with Temporal Mix of Experts

Transformer-based architectures dominate time series modeling by enabling global attention over all timestamps, yet their rigid 'one-size-fits-all' context aggregation fails to address two critical challenges in real-world data: (1) inherent lag effects, where the relevance of historical timestamps to a query varies dynamically; (2) anomalous segments, which introduce noisy signals that degrade forecasting accuracy. To resolve these problems, we propose the Temporal Mix of Experts (TMOE), a novel attention-level mechanism that reimagines key-value (K-V) pairs as local experts (each specialized in a distinct temporal context) and performs adaptive expert selection for each query via localized filtering of irrelevant timestamps. Complementing this local adaptation, a shared global expert preserves the Transformer's strength in capturing long-range dependencies. We then replace the vanilla attention mechanism in popular time-series Transformer frameworks (i.e., PatchTST and Timer) with TMOE, without extra structural modifications, yielding our specific version TimeExpert and general version TimeExpert-G. Extensive experiments on seven real-world long-term forecasting benchmarks demonstrate that TimeExpert and TimeExpert-G outperform state-of-the-art methods. Code is available at https://github.com/xwmaxwma/TimeExpert.

cs.LG

Multipoint distributions of the KPZ fixed point with compactly supported initial conditions

The KPZ fixed point is a universal limiting space-time random field for the Kardar-Parisi-Zhang universality class. While the joint law of the KPZ fixed point at a fixed time has been studied extensively, the multipoint distributions of the KPZ fixed point in the general space-time plane are much less well understood. More explicitly, formulas were only available for the narrow wedge initial condition arXiv:1906.01053, arXiv:1907.09876 and the flat initial condition arXiv:1907.09876 for the multipoint distributions, and the half-Brownian and Brownian initial conditions arXiv:2010.07357v1, arXiv:2504.19975 for the two-point distributions. In this paper, we obtain the first formula for the space-time joint distributions of the KPZ fixed point with general initial conditions of compact support. The formula is obtained through taking $1:2:3$ KPZ scaling limit of the multipoint distribution formulas for the totally asymmetric simple exclusion process (TASEP). A key ingredient is a probabilistic representation, inspired by arXiv:1701.00018, of the kernel encoding the initial condition for TASEP, which was first defined through an implicit characterization in arXiv:1907.09876. Moreover, we also verify that the equal time version of our formula matches the path integral formula in arXiv:1701.00018 for the KPZ fixed point when the initial condition is of compact support.

math.PR

Limiting one-point fluctuations of the geodesic in the directed landscape near the endpoints when the geodesic length goes to infinity

We consider the limiting fluctuations of the geodesic in the directed landscape, conditioning on its length going to infinity. It was shown in \cite{Liu22b,Ganguly-Hegde-Zhang23} that when the directed landscape $\mathcal{L}(0,0;0,1) = L$ becomes large, the geodesic from $(0,0)$ to $(0,1)$ lies in a strip of size $O(L^{-1/4})$ and behaves like a Brownian bridge if we zoom in the strip by a factor of $L^{1/4}$. Moreover, the length along the geodesic with respect to the directed landscape fluctuates of order $O(L^{1/4})$ and its limiting one-point distribution is Gaussian \cite{Liu22b}. In this paper, we further zoom in a smaller neighborhood of the endpoints when $\mathcal{L}(0,0;0,1) = L$ or $\mathcal{L}(0,0;0,1) \ge L$, and show that there is a critical scaling window $L^{-3/2}:L^{-1}:L^{-1/2}$ for the time, geodesic location, and geodesic length, respectively. Within this scaling window, we find a nontrivial limit of the one-point joint distribution of the geodesic location and length as $L\to\infty$. This limiting distribution, if we tune the time parameter to infinity, converges to the joint distribution of two independent Gaussian random variables, which is consistent with the results in \cite{Liu22b}. We also find a surprising connection between this limiting distribution and the one-point distribution of the upper tail field of the KPZ fixed point recently obtained in \cite{Liu-Zhang25}.

math.PR

On the multipoint distribution formulas of the parabolic Airy process

The parabolic Airy process is the Airy$_2$ process minus a parabola, initially defined by its finite-dimensional distributions, which are given by a Fredholm determinant formula with the extended Airy kernel. This process is also the one-time spatial marginal of the KPZ fixed point with the narrow wedge initial condition. There are two formulas for the space-time multipoint distribution of the KPZ fixed point with the narrow wedge initial condition obtained by arXiv:1906.01053 and arXiv:1907.09876. Especially, the equal-time case of arXiv:1907.09876 gives a different formula of the multipoint distribution of the parabolic Airy process. In this paper, we present a direct proof that this formula matches the one with the extended Airy kernel. Some byproducts in the proof include several new formulas for the parabolic Airy process, and a generalization of the Andreief's identity.

math.PR