SearcharxivSearch

arXiv subjects

Yingjie Zhou

Publications and source records attributed to Yingjie Zhou.

At least 19 recordsLinked to original sources

Femtoscopy as a New Probe of the Nuclear Equation of State

Femtoscopic correlations are widely regarded as precision probes of hadronic interactions through vacuum final-state interactions after kinetic freeze-out. Here we demonstrate that, in baryon-rich heavy-ion collisions, the nuclear mean field generates an additional dynamical contribution to femtoscopic correlations during the transport evolution. Using the Parton-Hadron-Quantum-Molecular Dynamics (PHQMD) transport approach, we investigate proton-proton, proton-$Λ$, three-proton, and proton-proton-$Λ$ correlations in Au+Au collisions at $\sqrt{s_{\rm NN}}=3$, 4.5, 7.7, and 19.6 GeV. We find that the nuclear mean field produces a characteristic low-$k^*$ enhancement that is strongest at the lowest beam energies and gradually disappears with increasing collision energy. Furthermore, both the stiffness and the momentum dependence of the nuclear equation of state leave distinct signatures in the femtoscopic correlation functions, with higher-order correlations exhibiting substantially enhanced sensitivity compared with conventional two-particle observables. Our results demonstrate that femtoscopy extends beyond its traditional role as a tool for studying hadronic interactions and serve as a new class of microscopic observables for the nuclear equation of state, complementary to collective flow and subthreshold strangeness production, thereby opening a new avenue for exploring dense baryonic matter in low-energy heavy-ion collisions.

nucl-th

Resolving the $ϕ$-meson directed-flow puzzle by multi-step meson--baryon dynamics

Recent STAR measurements at fixed-target Beam Energy Scan energies have revealed an unexpectedly large directed flow of $ϕ$ mesons in Au+Au collisions, comparable to that of protons and $Λ$ baryons and much stronger than that of light strange mesons. Since the $ϕ$ is a hidden-strangeness meson with relatively weak interactions with non-strange hadrons, this observation has been interpreted as a possible signal of unconventional baryonic dynamics or exotic baryonic resonances coupled to the $ϕ$ channel. Within the framework of the Parton-Hadron-Quantum-Molecular-Dynamics(PHQMD) model, we demonstrate that in the high baryon density region, $ϕ$ mesons are produced predominantly through multi-step meson--baryon and meson--hyperon reactions, whose transition amplitudes are constrained by a coupled-channel $T$-matrix calculation based on an extended SU(6) chiral effective Lagrangian. Together with the in-medium broadening of the $ϕ$ spectral function, these baryon-driven production channels enhance near-threshold $ϕ$ production and imprint the collective motion of the baryon-rich source on the produced $ϕ$ mesons.

nucl-th

Sequential Clusterization of Light Nuclei and Hypernuclei in Heavy-Ion Collisions within a Wigner Function Coalescence Framework

We investigate the formation of light nuclei and hypernuclei in Au+Au collisions at $\sqrt{s_{NN}}=3~\mathrm{GeV}$ within a coalescence framework embedded in the microscopic N-body Parton-Hadron-Quantum-Molecular Dynamics (PHQMD) transport model. The Wigner phase-space distributions employed in the coalescence calculation are constructed from realistic $N$-body wave functions obtained by solving the Schrödinger equation in the hyperspherical harmonics formalism, providing a solid and parameter-free description of nuclear clusters and hypernuclei. By comparing calculated rapidity distributions with STAR data, we extract species-dependent coalescence times, revealing a non-universal formation pattern among different clusters. The resulting yields and kinematic distributions of light nuclei and hypernuclei are systematically analyzed and shown to be sensitive to the underlying wave-function structure and formation time. In addition, we explore cluster-nucleon formation channels for $A=4$ systems. These additional channels improve the description of ${}^{4}\mathrm{He}$ and ${}^{4}_Λ\mathrm{H}$ yields and help address the underestimation of $A=4$ cluster production in theoretical approaches. Finally, we provide predictions for heavier hypernuclei, including ${}^{5}_Λ\mathrm{He}$ and ${}^{5}_{ΛΛ}\mathrm{He}$, which are of interest for future experimental measurements.

nucl-th

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, \emph{static monitoring} and \emph{dynamic reasoning}, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution traces, the curated benchmark corpus ProcessData contains \textasciitilde 58k question-answer pairs across 260 manipulation tasks, which is further split into ProcessData-SFT and ProcessData-Eval for post-training and evaluation purposes. Extensive evaluation of various VLMs on ProcessData-Eval reveals broad limitations across 12 diagnostic task families, suggesting current models still lack robust process-aware understanding of manipulation executions. But with ProcessData-SFT, the post-trained \textit{Qwen2.5-VL-7B} and \textit{InternVL-3-8B} exhibit consistent gains on local state, motion, progress, and primitive-aware cues. These results demonstrate that RoboProcessBench serves as both an evaluation benchmark and a learnable supervision source for developing VLMs capable of monitoring and evaluating robotic manipulation processes. Project webpage: \href{https://processbench-2026.github.io/RoboProcessBench-Web/}{https://processbench-2026.github.io}.

cs.RO

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary viewpoints become available. Building on this insight, we propose Reason, then Re-reason (ReRe), a training-free, inference-time framework with two phases: in the Reason Phase, an MLLM forms a spatial hypothesis from the original video; in the Re-reason Phase, it verifies or revises the hypothesis by observing a synthesized novel-view video. To enable effective cross-view revisiting, we design a Geometry-to-Video pipeline that renders strategically complementary novel views from predicted 3D geometry. These views feature an elevated, oblique perspective with scene-spanning coverage, while preserving the MLLM's native video interface without architectural modifications. Extensive evaluations on VSI-Bench and STI-Bench demonstrate that ReRe substantially boosts open-source MLLMs to rival proprietary state-of-the-art performance. Project page: https://zhenjiemao.github.io/ReRe/

cs.CV

An algebraic multiscale preconditioner for large sparse SPD matrices

We present a two-grid algebraic multiscale preconditioner for large sparse symmetric positive definite systems arising from elliptic problems with highly heterogeneous coefficients. The coarse space is constructed directly from the system matrix by graph partitioning and local generalized eigenvalue solvers, yielding basis functions that capture the low-energy modes responsible for slow convergence. The method requires no geometric information, making it suitable for unstructured and matrix-only settings, and its construction is naturally parallelizable. Numerical results for heterogeneous Darcy flow problems show robustness with respect to coefficient contrast and problem size, better performance than standard algebraic multigrid on challenging large-scale cases, and good parallel scalability.

math.NA

Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention

Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inherit the dense latent representations and quadratic-cost global modeling paradigm developed for high-resolution image synthesis, causing computation, memory usage, and inference latency to scale unfavorably with resolution and thus limiting practical deployment. We argue that the key bottleneck lies not in insufficient restoration priors, but in excessive token redundancy and costly token interactions during high-resolution restoration. Motivated by this observation, we revisit Real-ISR from the perspectives of compact latent representation and linear-complexity modeling, and propose SANA-SR, an efficient one-step restoration framework. Specifically, SANA-SR employs a deep compression autoencoder with a 32x compression ratio to drastically reduce latent tokens while preserving restoration-relevant structures and textures. On top of this compact latent space, we introduce a linear-attention DiT with LoRA fine-tuning, enabling efficient high-resolution restoration with linear-complexity token mixing. Extensive experiments on all benchmark datasets demonstrate that SANA-SR achieves highly competitive and often superior quantitative performance against existing methods, while restoring clearer and more realistic textures. Moreover, after pruning, the deployed model runs in 0.019s with 407.95G MACs and 344M parameters, highlighting its strong potential for practical mobile deployment.

cs.CV

Beyond Extrapolation: Knowledge Utilization Paradigm with Bidirectional Inspiration for Time Series Forecasting

Time-series forecasting is critical in various scenarios, such as energy, transportation, and public health. However, most existing forecasters rely primarily on one-way inference, \textit{i.e.}, mapping \textbf{history} to \textbf{target}, and overlook the structural information provided by a revised natural chain (``\textbf{history} (model input) -- \textbf{target} (ground-truth output) -- \textbf{post-target continuation}''). The post-target continuation records how trajectories evolve after the target, which can help stabilize forecasting, but it is not observable at inference time. In this work, we aim to obtain an approximate proxy of the post-target continuation for the current input, providing structural knowledge for bidirectional forecasting. This idea is instantiated as KUP-BI (Knowledge Utilization Paradigm with Bidirectional Inspiration), a new time-series modeling paradigm that distills continuation-style knowledge (as an approximate post-target continuation proxy) from a \emph{train-only} historical library and integrates it into standard forecasting backbones. The input stream and the continuation-proxy stream are fused via a lightweight feature-level gating module. This design does not introduce information beyond what is already contained in the training trajectories; instead, it provides a structured inductive bias that helps backbones exploit typical continuation patterns rather than relying solely on parametric extrapolation. Experimental results on six public datasets show that KUP-BI consistently improves the forecasting performance of state-of-the-art models, with small additional overhead.

cs.LG

SeesawNet: Towards Non-stationary Time Series Forecasting with Balanced Modeling of Common and Specific Dependencies

Instance normalization (IN) is widely used in non-stationary multivariate time series forecasting to reduce distribution shifts and highlight common patterns across samples. However, IN can over-smooth instance-specific structural information that is essential for modeling temporal and cross-channel heterogeneity. While prior methods further suppress distribution discrepancies or attempt to recover temporal specific dependencies, they often ignore a central tension: how to adaptively model common and instance-specific dependency based on each instance's non-stationary structures. To address this dilemma, we propose SeesawNet, a unified architecture that dynamically balances common and instance-specific dependency modeling in both temporal and channel dimensions. At its core is Adaptive Stationary-Nonstationary Attention (ASNA), which captures common dependencies from normalized sequences and specific dependencies from raw sequences, and adaptively fuses them according to instance-level non-stationarity. Built upon ASNA, SeesawNet alternates dedicated temporal and channel relationship modeling to jointly capture long-range and cross-variable dependencies. Extensive experiments on multiple real-world benchmarks demonstrate that SeesawNet consistently outperforms state-of-the-art methods.

cs.LG

Learning Feature Encoder with Synthetic Anomalies for Weakly Supervised Graph Anomaly Detection

Weakly supervised graph anomaly detection aims to unveil unusual graph instances, e.g., nodes, whose behaviors significantly differ from normal ones, given only a limited number of annotated anomalies and abundant unlabeled samples. A major challenge is to learn a meaningful latent feature representation that reduces intra-class variance among normal data while remaining highly sensitive to anomalies. Although recent works have applied self-supervised feature learning for graph anomaly detection, their strategies are not specifically tailored to its unique requirements, motivating our exploration of a more domain-specific approach. In this paper, we introduce a weakly supervised graph anomaly detection method that leverages a feature learning strategy tailored for graph anomalies. Our approach is built upon a multi-task learning scheme that extracts robust feature representations through synthesized anomalies. We generate synthetic anomalies by perturbing the normal graph in various ways and assign a dedicated detection head to each anomaly type, ensuring that learned features are sensitive to potential deviations from normal patterns. Although synthetic anomalies may not perfectly replicate real-world patterns, they provide valuable auxiliary data for effective feature learnin, much like features learned from ImageNet classification transfer to downstream vision tasks. Additionally, we adopt a two-phase learning strategy: an initial warm-up phase using only synthetic samples, followed by a full-training phase integrating both tasks, to balance the influence of synthetic and real data. Extensive experiments on public datasets demonstrate the superior performance of our method over its competitors. Code is available at https://github.com/yj-zhou/SAWGAD.

cs.LG

Efficient Multiscale Methods for Highly Heterogeneous Spatial Network Models

Modeling complex spatial networks with multiscale heterogeneity poses significant mathematical and computational challenges. Lacking explicit PDE discretizations and facing excessive degrees of freedom, conventional methods often become computationally prohibitive. To address these challenges, we propose an efficient multiscale modeling for highly heterogeneous spatial networks. We construct multiscale basis functions tailored to spatial network models with heterogeneous edge weights and node degrees. A key novelty is that the proposed method doesn't introduce geometric parameters (such as Dirichlet nodes, distances, or mesh sizes), thereby preserving its purely algebraic nature and ensuring broad applicability. By incorporating a subgraph-wise estimate, we define a Poincaré constant $C_{\mathrm{po}}$ that renders the method independent of the underlying graph geometry. Then through an appropriate choice of the number of graph oversampling layers, we establish an $O(C_{\mathrm{po}})$ convergence independent of the local heterogeneity contrast. Notably, our scheme operates entirely within an algebraic framework, eliminating the need for Dirichlet nodes and positive-definiteness on specific matrices arising in the model. This flexibility enables the simulation of a wider range of physical models and accommodates various boundary conditions. Rigorous theoretical analyses are provided under suitable assumptions, and extensive numerical experiments validate the effectiveness of the proposed approach.

math.NA

Q-Agent: Quality-Driven Chain-of-Thought Image Restoration Agent through Robust Multimodal Large Language Model

Image restoration (IR) often faces various complex and unknown degradations in real-world scenarios, such as noise, blurring, compression artifacts, and low resolution, etc. Training specific models for specific degradation may lead to poor generalization. To handle multiple degradations simultaneously, All-in-One models might sacrifice performance on certain types of degradation and still struggle with unseen degradations during training. Existing IR agents rely on multimodal large language models (MLLM) and a time-consuming rolling-back selection strategy neglecting image quality. As a result, they may misinterpret degradations and have high time and computational costs to conduct unnecessary IR tasks with redundant order. To address these, we propose a Quality-Driven agent (Q-Agent) via Chain-of-Thought (CoT) restoration. Specifically, our Q-Agent consists of robust degradation perception and quality-driven greedy restoration. The former module first fine-tunes MLLM, and uses CoT to decompose multi-degradation perception into single-degradation perception tasks to enhance the perception of MLLMs. The latter employs objective image quality assessment (IQA) metrics to determine the optimal restoration sequence and execute the corresponding restoration algorithms. Experimental results demonstrate that our Q-Agent achieves superior IR performance compared to existing All-in-One models.

eess.IV

A Data-Guided Coalescence Model for Light Nuclei and Hypernuclei Production in Relativistic Heavy-Ion Collisions at $\sqrt{s_{\rm{NN}}} = 3$--200 GeV

The production of light hypernuclei in relativistic heavy-ion collisions provides a unique opportunity to probe hyperon--nucleon interactions and possible three-body forces, which are central to the resolution of the hyperon puzzle in neutron star matter. In this work, we develop a data-guided coalescence framework in which the source size is extracted from proton and deuteron yields and used, together with measured proton and $Λ$ spectra, to predict the production of $A=3$ nuclei $(t,{}^{3}\rm{He})$ and hypernuclei $({}^{3}_Λ\rm{H})$ in $\sqrt{s_{\rm{NN}}}=3$--200~GeV Au+Au collisions. For tritons, calculations with a Gaussian wave function reproduce experimental spectra and yield ratios across this energy range. For the hypertriton, the model predictions are highly sensitive to the assumed wave function. This sensitivity is strongest at low collision energies and in low-multiplicity environments, implying that such conditions are particularly valuable for probing hypernuclear structure.

nucl-th

Surveillance Facial Image Quality Assessment: A Multi-dimensional Dataset and Lightweight Model

Surveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications.

eess.IV

Distributed Online Economic Dispatch with Time-Varying Coupled Inequality Constraints

We investigate the distributed online economic dispatch problem for power systems with time-varying coupled inequality constraints. The problem is formulated as a distributed online optimization problem in a multi-agent system. At each time step, each agent only observes its own instantaneous objective function and local inequality constraints; agents make decisions online and cooperate to minimize the sum of the time-varying objectives while satisfying the global coupled constraints. To solve the problem, we propose an algorithm based on the primal-dual approach combined with constraint-tracking. Under appropriate assumptions that the objective and constraint functions are convex, their gradients are uniformly bounded, and the path length of the optimal solution sequence grows sublinearly, we analyze theoretical properties of the proposed algorithm and prove that both the dynamic regret and the constraint violation are sublinear with time horizon T. Finally, we evaluate the proposed algorithm on a time-varying economic dispatch problem in power systems using both synthetic data and Australian Energy Market data. The results demonstrate that the proposed algorithm performs effectively in terms of tracking performance, constraint satisfaction, and adaptation to time-varying disturbances, thereby providing a practical and theoretically well-supported solution for real-time distributed economic dispatch.

math.OC

Probing of EoS with clusters and hypernuclei

The study of the nuclear equation-of-state (EoS) is a one of the primary goals of experimental and theoretical heavy-ion physics. The comparison of recent high statistics data from the STAR Collaboration with transport models provides a unique possibility to address this topic in a yet unexplored energy domain. Employing the microscopic N-body Parton-Hadron-Quantum-Molecular Dynamics (PHQMD) transport approach, which allows to describe the propagation and interactions of hadronic and partonic degrees of freedom including cluster and hyper-nucleus formation and dynamics, we investigate the influence of different EoS on bulk observables, the multiplicity, $p_T$ and rapidity distributions of protons, $Λ$s and clusters up to A=4 as well as their influence on the collective flow. We explore three different EoS: two static EoS, dubbed 'soft' and 'hard', which differ in the compressibility modulus, as well as a soft momentum dependent EoS. We find that a soft momentum dependent EoS reproduces most baryon and cluster observables, including the flow observables, quantitatively, however, hard EOS show a similar trend.

nucl-th

EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans

Speech-driven Talking Human (TH) generation, commonly known as "Talker," currently faces limitations in multi-subject driving capabilities. Extending this paradigm to "Multi-Talker," capable of animating multiple subjects simultaneously, introduces richer interactivity and stronger immersion in audiovisual communication. However, current Multi-Talkers still exhibit noticeable quality degradation caused by technical limitations, resulting in suboptimal user experiences. To address this challenge, we construct THQA-MT, the first large-scale Multi-Talker-generated Talking Human Quality Assessment dataset, consisting of 5,492 Multi-Talker-generated THs (MTHs) from 15 representative Multi-Talkers using 400 real portraits collected online. Through subjective experiments, we analyze perceptual discrepancies among different Multi-Talkers and identify 12 common types of distortion. Furthermore, we introduce EvalTalker, a novel TH quality assessment framework. This framework possesses the ability to perceive global quality, human characteristics, and identity consistency, while integrating Qwen-Sync to perceive multimodal synchrony. Experimental results demonstrate that EvalTalker achieves superior correlation with subjective scores, providing a robust foundation for future research on high-quality Multi-Talker generation and evaluation.

cs.CV

DLADiff: A Dual-Layer Defense Framework against Fine-Tuning and Zero-Shot Customization of Diffusion Models

With the rapid advancement of diffusion models, a variety of fine-tuning methods have been developed, enabling high-fidelity image generation with high similarity to the target content using only 3 to 5 training images. More recently, zero-shot generation methods have emerged, capable of producing highly realistic outputs from a single reference image without altering model weights. However, technological advancements have also introduced significant risks to facial privacy. Malicious actors can exploit diffusion model customization with just a few or even one image of a person to create synthetic identities nearly identical to the original identity. Although research has begun to focus on defending against diffusion model customization, most existing defense methods target fine-tuning approaches and neglect zero-shot generation defenses. To address this issue, this paper proposes Dual-Layer Anti-Diffusion (DLADiff) to defense both fine-tuning methods and zero-shot methods. DLADiff contains a dual-layer protective mechanism. The first layer provides effective protection against unauthorized fine-tuning by leveraging the proposed Dual-Surrogate Models (DSUR) mechanism and Alternating Dynamic Fine-Tuning (ADFT), which integrates adversarial training with the prior knowledge derived from pre-fine-tuned models. The second layer, though simple in design, demonstrates strong effectiveness in preventing image generation through zero-shot methods. Extensive experimental results demonstrate that our method significantly outperforms existing approaches in defending against fine-tuning of diffusion models and achieves unprecedented performance in protecting against zero-shot generation.

eess.IV