SearcharxivSearch

arXiv subjects

Zhiwen Wang

Publications and source records attributed to Zhiwen Wang.

At least 19 recordsLinked to original sources

Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.

cs.CV

Learning to Rank for Selected Configuration Interaction

The accurate description of electron correlation is a central challenge in computational chemistry, with selected configuration interaction (SCI) emerging as a powerful tool to approach the full CI limit. While recent machine learning (ML) integrations have accelerated determinant selection, existing regression and classification approaches suffer from a fundamental objective-loss mismatch: they evaluate the importance of determinants in isolation without explicitly accounting for their relative importance ranking. Here, we introduce ranking configuration interaction (RCI), a novel ML-supported SCI framework that reframes determinant selection as a pairwise ranking problem. Building upon a Transformer-based architecture to capture complex, non-local orbital dependencies, RCI progressively optimizes the partial ordering of determinants. By doing so, RCI aligns the training objective more closely with the intrinsic ranking nature of SCI. Extensive benchmarks across both plane-wave and Gaussian basis sets, including the molecules N$_2$, CO, H$_2$O, NH$_3$, and C$_2$, demonstrate the efficiency of RCI. Compared to previously reported classification baselines, RCI consistently accelerates convergence-reducing overall computational time by 23% to over 50% depending on the system, and requiring only 55% of the determinant count in representative cases such as N$_2$ and CO. Furthermore, RCI exhibits robust performance and reaches chemical accuracy on the highly challenging iron-sulfur cluster using only 12% of the full CI space. Notably, RCI outperforms recent regression-based SCI methods by delivering a more than 15% improvement in accuracy at comparable determinant counts. RCI also demonstrates higher efficiency than heat-bath CI on the strongly correlated chromium dimer, yielding a compact and accurate wavefunction.

physics.chem-ph

Where Do CoT Training Gains Land in LLM based Agents?

Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning. We therefore ask what CoT training is actually improving: is the model getting better at changing its action through generated reasoning, or is it getting better at predicting the action directly from the prompt? We study this question by comparing \emph{prompt actions} (predicting action without CoT) with CoT actions (predicting action with CoT). Across checkpoints, prompt-action quality improves substantially. While interacting with the environment, the relative advantage of CoT actions over prompt actions remains similar, showing that CoT training does not widen the advantage of CoT reasoning, and it helps to improve the quality of prompt actions. We further find that later checkpoints are less likely to revise the action in response to CoT, suggesting greater reliance on the prompt. Motivated by these patterns, we selectively mask action-token supervision on a fraction of training examples. This intervention improves out-of-domain generalization.

cs.AI

Extremal results on the second largest eigenvalue of graphs with given order

In this paper, we demonstrate the effects on the second largest eigenvalue $λ_2(G)$ of a connected graph $G$ after edge addition or deletion. In 1989, Chung, Graham and Wilson showed $\max\{|λ_2|,|λ_n|\}>Ω(n)$ for dense $K_{r+1}$-free graphs of order $n$, giving spectral comprehension of existence of large clique or independent set, respect to Ramsey theory. Applying the results of effects on $λ_2$ after edge operations, we determine the maximum value of $λ_2$ among all $K_{r+1}$-free connected graphs with given order, and completely characterize the extremal graphs. Moreover, for arbitrary given graph $F$, we investigates the maximum second largest $λ_2(G)$ among $F$-free connected graphs of order $n$. Let $ρ^*(n,F)$ be the maximum spectral radius of $F$-free graphs on $n\ge n_F$ vertices, and $G^*(n,F)$ be a graph with its spectral radius $ρ\big(G^*(n,F)\big)=ρ^*(n,F)$. We prove that, for an $F$-free connected graph $G$ of order $n\ge f(n_F)$, \\(1) if $n$ is odd, then $$λ_2(G)\leρ^*\left(\frac{n-1}{2},F\right)$$ with equality if and only if $G\in \mathcal{I}\big(G^*(\frac{n-1}{2},F),G^*(\frac{n-1}{2},F)\big)$; and\\ (2) if $n$ is even, and $F$ does not contain cut edges, then the graph $G^†$ with the maximum second largest eigenvalue satisfies $$λ_2(G^†)=ρ^*\left(\frac{n}{2},F\right)-o(1)$$ and $G^†\in \mathcal{E}\big(H_1,H_2\big)$, where $H_1$ and $H_2$ are $F$-saturated graphs on $\frac{n}{2}$ vertices. In particular, other than a complete graph $K_{r+1}$, when $F$ is a book graph $B_{k+1}$ or an odd cycle $C_{2k+1}$, we are able to determine the maximum second largest eigenvalue for $F$-free connected graphs of given order, and completely characterize the extremal graphs.

math.CO

Upper bounds of the second largest eigenvalue of graphs

Let $λ_i(G)$ denote the $i$-th largest eigenvalue of adjacency matrix of a graph $G$. Gerschgorin's Theorem indicates $λ_1(G)$ belongs to the largest disk, i.e., $λ_1(G)\leΔ_1(G)$, where $Δ_i(G)$ is the $i$-th largest degree of $G$. We show that $λ_2(G)$ lies in the second largest disk. That is, in detail, $$λ_2(G)<Δ_2(G)-\frac{1}{n^2}.$$ A classical theorem proved by Hong [\textit{Linear Algebra Appl.} 1988] states that $λ_1(G)\le\sqrt{2m-n+1}$ for a connected graph $G$ with $n$ vertices and $m$ edges, where the equality holds if and only if $G$ is a star $S_n$ or a complete graph $K_n$. We give a refinement of Hong's theorem by showing $$λ_1(G)<\sqrt{2m-n}$$ for any connected graph $G\not\in\left\{S_n,S^1_{n-1},K_n,K^1_{n-1}\right\}$. Based on this improved upper bound of $λ_1(G)$, for a connected graph $G$ with $n$ vertices and $m$ edges, we are able to prove a sharp upper bound of $λ_2(G)$ that $$λ_2(G)\le\sqrt{m-\frac{n}{2}-\frac{1}{2}},$$ except $G$ is obtained from two disjoint $S_\frac{n}{2}$ by adding an edge between a pendant vertex of each star. Moreover, we provide a complete characterization to extremal graphs attaining the equality.

math.CO

GeoSSA: Geometric Sparrow Search Algorithm for UAV Path Planning and Engineering Design Optimization

The Sparrow Search Algorithm (SSA), characterized by its simple structure and ease of implementation, nevertheless suffers from an insufficient balance between exploration and exploitation, making it prone to premature convergence and slow optimization progress. To address these shortcomings, this paper proposes a Geometric Sparrow Search Algorithm (GeoSSA). By integrating Good Nodes Set initialization, a Sine-Cosine Enhanced Producer position update strategy, and a Triangular-Walk Enhanced Edge Sparrow update strategy, GeoSSA significantly improves the global exploration ability, local exploitation efficiency, and convergence stability of the original SSA. To thoroughly validate the effectiveness of GeoSSA, we conducted ablation studies, qualitative analysis, and comparative experiments on 23 benchmark functions against state-of-the-art algorithms. Experimental results show that GeoSSA achieves the best or near-best performance in terms of average fitness, standard deviation, Wilcoxon tests, and Friedman rankings, with an Overall Effectiveness ($OE$) of 95.65\%. Its overall performance is significantly superior to all compared algorithms. In three-dimensional UAV path planning tasks, GeoSSA demonstrates excellent stability and superior path quality. In four categories of engineering design optimization problems, GeoSSA consistently attains the highest solution accuracy and strongest stability. GeoSSA not only exhibits outstanding global optimization performance on standard benchmark functions but also shows strong robustness and generalization ability in practical applications such as UAV path planning and engineering design. Therefore, GeoSSA provides an efficient and reliable solution framework for complex optimization problems.

cs.CE

Bandwidth Selection of Density Estimators over Treespaces

A kernel density estimator (KDE) is one of the most popular non-parametric density estimators. In this paper we focus on a best bandwidth selection method for use in an analogue of a classical KDE using the tropical symmetric distance, known as a tropical KDE, for use over the space of phylogenetic trees. We propose the likelihood cross validation (LCV) for selecting the bandwidth parameter for the KDE over the space of phylogenetic trees. In this paper, first, we show the explicit optimal solution of the best-fit bandwidth parameter via the LCV for tropical KDE over the space of phylogenetic trees. Then, computational experiments with simulated datasets generated under the multi-species coalescent (MSC) model show that a tropical KDE with the best-fit bandwidth parameter via the LCV perform better than a tropical KDE with an estimated best-fit bandwidth parameter via nearest neighbors in terms of accuracy and computational time. Lastly, we apply our method an empirical data from the Apicomplexa genome.

q-bio.PE

NAWOA-XGBoost: A Novel Model for Early Prediction of Academic Potential in Computer Science Students

Whale Optimization Algorithm (WOA) suffers from limited global search ability, slow convergence, and tendency to fall into local optima, restricting its effectiveness in hyperparameter optimization for machine learning models. To address these issues, this study proposes a Nonlinear Adaptive Whale Optimization Algorithm (NAWOA), which integrates strategies such as Good Nodes Set initialization, Leader-Followers Foraging, Dynamic Encircling Prey, Triangular Hunting, and a nonlinear convergence factor to enhance exploration, exploitation, and convergence stability. Experiments on 23 benchmark functions demonstrate NAWOA's superior optimization capability and robustness. Based on this optimizer, an NAWOA-XGBoost model was developed to predict academic potential using data from 495 Computer Science undergraduates at Macao Polytechnic University (2009-2019). Results show that NAWOA-XGBoost outperforms traditional XGBoost and WOA-XGBoost across key metrics, including Accuracy (0.8148), Macro F1 (0.8101), AUC (0.8932), and G-Mean (0.8172), demonstrating strong adaptability on multi-class imbalanced datasets.

cs.CE

Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images

Text-based pedestrian search (TBPS) in full images aims to locate a target pedestrian in untrimmed images using natural language descriptions. However, in complex scenes with multiple pedestrians, existing methods are limited by uncertainties in detection and matching, leading to degraded performance. To address this, we propose UPD-TBPS, a novel framework comprising three modules: Multi-granularity Uncertainty Estimation (MUE), Prototype-based Uncertainty Decoupling (PUD), and Cross-modal Re-identification (ReID). MUE conducts multi-granularity queries to identify potential targets and assigns confidence scores to reduce early-stage uncertainty. PUD leverages visual context decoupling and prototype mining to extract features of the target pedestrian described in the query. It separates and learns pedestrian prototype representations at both the coarse-grained cluster level and the fine-grained individual level, thereby reducing matching uncertainty. ReID evaluates candidates with varying confidence levels, improving detection and retrieval accuracy. Experiments on CUHK-SYSU-TBPS and PRW-TBPS datasets validate the effectiveness of our framework.

cs.CV

Patient-Level Anatomy Meets Scanning-Level Physics: Personalized Federated Low-Dose CT Denoising Empowered by Large Language Model

Reducing radiation doses benefits patients, however, the resultant low-dose computed tomography (LDCT) images often suffer from clinically unacceptable noise and artifacts. While deep learning (DL) shows promise in LDCT reconstruction, it requires large-scale data collection from multiple clients, raising privacy concerns. Federated learning (FL) has been introduced to address these privacy concerns; however, current methods are typically tailored to specific scanning protocols, which limits their generalizability and makes them less effective for unseen protocols. To address these issues, we propose SCAN-PhysFed, a novel SCanning- and ANatomy-level personalized Physics-Driven Federated learning paradigm for LDCT reconstruction. Since the noise distribution in LDCT data is closely tied to scanning protocols and anatomical structures being scanned, we design a dual-level physics-informed way to address these challenges. Specifically, we incorporate physical and anatomical prompts into our physics-informed hypernetworks to capture scanning- and anatomy-specific information, enabling dual-level physics-driven personalization of imaging features. These prompts are derived from the scanning protocol and the radiology report generated by a medical large language model (MLLM), respectively. Subsequently, client-specific decoders project these dual-level personalized imaging features back into the image domain. Besides, to tackle the challenge of unseen data, we introduce a novel protocol vector-quantization strategy (PVQS), which ensures consistent performance across new clients by quantifying the unseen scanning code as one of the codes in the scanning codebook. Extensive experimental results demonstrate the superior performance of SCAN-PhysFed on public datasets.

eess.IV

The Transition from Galaxy-wide Gas Inflow to Outflow in Quasar Host Galaxies

Galactic-wide outflows driven by active galactic nuclei (AGNs) is a routinely invoked feedback mechanism in galaxy evolution models. Hitherto, the interplay among the interstellar gas on galactic scales, the propagation of AGN outflows and the fundamental AGN parameters during evolution remains elusive. Powerful nuclear outflows are found to favorably exist at early AGN stages usually associated with high accretion rates and weak narrow emission lines. In a sample of quasars emitting Mg II narrow absorption lines (NALs) from the Sloan Digital Sky Survey, we discover an unprecedented phenomenon where galaxy-scale inflow-dominated transforming into outflow-dominated gas accompanied by an increasing strength of the narrow [O III] line, at a confidence level of 6.7σ. The fact that nuclear outflows diminish while galaxy-wide outflows intensifies as AGNs evolve implies that early-stage outflows interact with interstellar medium on galactic scales and trigger the gradual transformation into galaxy-wide outflows, providing observational links to the hypothetical multi-stage propagation of AGN outflows that globally regulates galaxy evolution.

astro-ph.GA

JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image Enhancement

Low-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach.

cs.CV

Plaintext-Free Deep Learning for Privacy-Preserving Medical Image Analysis via Frequency Information Embedding

In the fast-evolving field of medical image analysis, Deep Learning (DL)-based methods have achieved tremendous success. However, these methods require plaintext data for training and inference stages, raising privacy concerns, especially in the sensitive area of medical data. To tackle these concerns, this paper proposes a novel framework that uses surrogate images for analysis, eliminating the need for plaintext images. This approach is called Frequency-domain Exchange Style Fusion (FESF). The framework includes two main components: Image Hidden Module (IHM) and Image Quality Enhancement Module~(IQEM). The~IHM performs in the frequency domain, blending the features of plaintext medical images into host medical images, and then combines this with IQEM to improve and create surrogate images effectively. During the diagnostic model training process, only surrogate images are used, enabling anonymous analysis without any plaintext data during both training and inference stages. Extensive evaluations demonstrate that our framework effectively preserves the privacy of medical images and maintains diagnostic accuracy of DL models at a relatively high level, proving its effectiveness across various datasets and DL-based models.

cs.CR

Privacy-Preserving Encrypted Low-Dose CT Denoising

Deep learning (DL) has made significant advancements in tomographic imaging, particularly in low-dose computed tomography (LDCT) denoising. A recent trend involves servers training powerful models with large amounts of self-collected private data and providing application programming interfaces (APIs) for users, such as Chat-GPT. To avoid model leakage, users are required to upload their data to the server model, but this way raises public concerns about the potential risk of privacy disclosure, especially for medical data. Hence, to alleviate related concerns, in this paper, we propose to directly denoise LDCT in the encrypted domain to achieve privacy-preserving cloud services without exposing private data to the server. To this end, we employ homomorphic encryption to encrypt private LDCT data, which is then transferred to the server model trained with plaintext LDCT for further denoising. However, since traditional operations, such as convolution and linear transformation, in DL methods cannot be directly used in the encrypted domain, we transform the fundamental mathematic operations in the plaintext domain into the operations in the encrypted domain. In addition, we present two interactive frameworks for linear and nonlinear models in this paper, both of which can achieve lossless operating. In this way, the proposed methods can achieve two merits, the data privacy is well protected and the server model is free from the risk of model leakage. Moreover, we provide theoretical proof to validate the lossless property of our framework. Finally, experiments were conducted to demonstrate that the transferred contents are well protected and cannot be reconstructed. The code will be released once the paper is accepted.

cs.CR

Solving multiscale elliptic problems by sparse radial basis function neural networks

Machine learning has been successfully applied to various fields of scientific computing in recent years. In this work, we propose a sparse radial basis function neural network method to solve elliptic partial differential equations (PDEs) with multiscale coefficients. Inspired by the deep mixed residual method, we rewrite the second-order problem into a first-order system and employ multiple radial basis function neural networks (RBFNNs) to approximate unknown functions in the system. To aviod the overfitting due to the simplicity of RBFNN, an additional regularization is introduced in the loss function. Thus the loss function contains two parts: the $L_2$ loss for the residual of the first-order system and boundary conditions, and the $\ell_1$ regularization term for the weights of radial basis functions (RBFs). An algorithm for optimizing the specific loss function is introduced to accelerate the training process. The accuracy and effectiveness of the proposed method are demonstrated through a collection of multiscale problems with scale separation, discontinuity and multiple scales from one to three dimensions. Notably, the $\ell_1$ regularization can achieve the goal of representing the solution by fewer RBFs. As a consequence, the total number of RBFs scales like $\mathcal{O}(\varepsilon^{-nτ})$, where $\varepsilon$ is the smallest scale, $n$ is the dimensionality, and $τ$ is typically smaller than $1$. It is worth mentioning that the proposed method not only has the numerical convergence and thus provides a reliable numerical solution in three dimensions when a classical method is typically not affordable, but also outperforms most other available machine learning methods in terms of accuracy and robustness.

math.NA

EVIL: Evidential Inference Learning for Trustworthy Semi-supervised Medical Image Segmentation

Recently, uncertainty-aware methods have attracted increasing attention in semi-supervised medical image segmentation. However, current methods usually suffer from the drawback that it is difficult to balance the computational cost, estimation accuracy, and theoretical support in a unified framework. To alleviate this problem, we introduce the Dempster-Shafer Theory of Evidence (DST) into semi-supervised medical image segmentation, dubbed Evidential Inference Learning (EVIL). EVIL provides a theoretically guaranteed solution to infer accurate uncertainty quantification in a single forward pass. Trustworthy pseudo labels on unlabeled data are generated after uncertainty estimation. The recently proposed consistency regularization-based training paradigm is adopted in our framework, which enforces the consistency on the perturbed predictions to enhance the generalization with few labeled data. Experimental results show that EVIL achieves competitive performance in comparison with several state-of-the-art methods on the public dataset.

cs.CV

The multiplicity of a Hermitian eigenvalue on graphs

For a graph $G$, let $\mathcal{S}(G)$ be the set consisting of Hermitian matrices whose graph is $G$. Denoted by $m_B(G,λ)$ the multiplicity of an eigenvalue $λ$ of $B(G)\in \mathcal{S}(G)$, we show that $m_B(G,λ)\le 2θ(G)+p(G)$ where $θ(G)$ and $p(G)$ are the cyclomatic number and the number of pendent vertices of $G$ respectively, and characterize the graphs attaining the equality. This is a generalization of a result on adjacency matrix by Wang et al.\cite{Wang1}. Moreover, they arose an open problem in \cite{Wang1}: \textit{characterize all graphs with $m_A(G,λ)=2θ(G)+p(G)-1$ for any eigenvalue $λ$ of its adjacency matrix.} In this paper, we completely characterize the graphs with $m_B(G,λ)=2θ(G)+p(G)-1$ for any eigenvalue $λ$ of an arbitrary Hermitian matrix $B(G)\in \mathcal{S}(G)$. This result provides a stronger answer to the above problem, and encompasses some previous known works considering $λ=-1$ or $0$ on the problem.

math.CO

Target-Aware Tracking with Long-term Context Attention

Most deep trackers still follow the guidance of the siamese paradigms and use a template that contains only the target without any contextual information, which makes it difficult for the tracker to cope with large appearance changes, rapid target movement, and attraction from similar objects. To alleviate the above problem, we propose a long-term context attention (LCA) module that can perform extensive information fusion on the target and its context from long-term frames, and calculate the target correlation while enhancing target features. The complete contextual information contains the location of the target as well as the state around the target. LCA uses the target state from the previous frame to exclude the interference of similar objects and complex backgrounds, thus accurately locating the target and enabling the tracker to obtain higher robustness and regression accuracy. By embedding the LCA module in Transformer, we build a powerful online tracker with a target-aware backbone, termed as TATrack. In addition, we propose a dynamic online update algorithm based on the classification confidence of historical information without additional calculation burden. Our tracker achieves state-of-the-art performance on multiple benchmarks, with 71.1\% AUC, 89.3\% NP, and 73.0\% AO on LaSOT, TrackingNet, and GOT-10k. The code and trained models are available on https://github.com/hekaijie123/TATrack.

cs.CV