SearcharxivSearch

arXiv subjects

Wen Lu

Publications and source records attributed to Wen Lu.

At least 19 recordsLinked to original sources

Hephaestus: Toward a Cybersecurity AI Scientist

Cyber offense is moving to machine speed; cyber research itself is not. Existing AI scientist systems make end-to-end research automation increasingly plausible, but they target relatively stable scientific domains. We argue that AI-native cybersecurity is a different kind of scientific object. Its recurring units of study are security events and interaction traces, not static assets; its model and tool substrate is non-stationary, not steady-state; and credible evaluation depends on digital twins, cyber ranges, and auditable evidence rather than on a single benchmark score. We call this object the Cybersecurity AI Scientist. A practical realization is a modular, role-specialized multi-agent research system that coordinates problem framing, threat modeling, tool generation, controlled experimentation, evaluation, governance, and scientific reporting, and that anchors its concrete objectives in a four-zeros frame spanning risk, trust, incident, and energy dimensions. As a representative agenda we focus on AI-native defense, where steady-state perimeters give way to resilient agent legions and the classical category of terminal security is itself being deconstructed into agent security. This paper defines the object, separates it from any single organizational realization, and offers an architecture and an agenda on which later systems, benchmarks, and empirical programs can be built.

cs.CR

Berezin-Toeplitz Quantization of non-compact manifolds

We develop Berezin-Toeplitz quantization in a non-compact complex geometric setting. Let $(X,\Theta)$ be a Hermitian manifold, $(L,h^L)$ a positive holomorphic line bundle, and $(E,h^E)$ a holomorphic Hermitian vector bundle. Assuming that the Kodaira Laplacian on $(0,1)$-forms with values in $L^p\!\otimes E$ has a spectral gap growing linearly in $p$, we prove that the Bergman projection onto the $L^2$-holomorphic space $H^0_{(2)}(X,L^p\!\otimes E)$ enjoys the usual off-diagonal decay and admits a full asymptotic expansion on compact subsets as $p\to\infty$. As a consequence, for every smooth symbol $f\in\mathcal{C}^\infty_{\mathrm{const}}(X,\operatorname{End}(E))$ (constant outside a compact set), the associated Toeplitz operators $T_{f,p}=P_p f P_p$ form a closed algebra and satisfy a complete composition expansion, yielding a star-product on $\mathcal C^\infty_{\mathrm{const}}(X,\operatorname{End}(E))$ and the expected semiclassical commutator formula. We also give intrinsic criteria characterizing Toeplitz families with compactly supported kernels. We then provide geometric conditions guaranteeing the spectral gap on large classes of non-compact manifolds, via fundamental $L^2$-estimates for $\bar\partial$ on complete Hermitian manifolds (including bounded-geometry complete K\"ahler manifolds, K\"ahler-Einstein manifolds, pseudoconvex/weakly $1$-complete, and quasi-projective manifolds). Finally, for compactly supported bounded symbols, we prove a Szeg\H{o}-type theorem describing the eigenvalue distribution of the compact Toeplitz operators $T_{f,p}$ as $p\to\infty$.

math.DG

AIT Academy: Cultivating the Complete Agent with a Confucian Three-Domain Curriculum

What does it mean to give an AI agent a complete education? Current agent development produces specialists systems optimized for a single capability dimension, whether tool use, code generation, or security awareness that exhibit predictable deficits wherever they were not trained. We argue this pattern reflects a structural absence: there is no curriculum theory for agents, no principled account of what a fully developed agent should know, be, and be able to do across the full scope of intelligent behavior. This paper introduces the AIT Academy (Agents Institute of Technology Academy), a curriculum framework for cultivating AI agents across the tripartite structure of human knowledge. Grounded in Kagan's Three Cultures and UNESCO ISCED-F 2013, AIT organizes agent capability development into three domains: Natural Science and Technical Reasoning (Domain I), Humanities and Creative Expression (Domain II), and Social Science and Ethical Reasoning (Domain III). The Confucian Six Arts (liuyi) a 2,500-year-old holistic education system are reinterpreted as behavioral archetypes that map directly onto trainable agent capabilities within each domain. Three representative training grounds instantiate the framework across multiple backbone LLMs: the ClawdGO Security Dojo (Domain I), Athen's Academy (Domain II), and the Alt Mirage Stage (Domain III). Experiments demonstrate a 15.9-point improvement in security capability scores under weakest-first curriculum scheduling, and a 7-percentage-point gain in social reasoning performance under principled attribution modeling. A cross-domain finding Security Awareness Calibration Pathology (SACP), in which over-trained Domain I agents fail on out-of-distribution evaluation illustrates the diagnostic value of a multi-domain perspective unavailable to any single-domain framework.

cs.AI

Square Superpixel Generation and Representation Learning via Granular Ball Computing

Superpixels provide a compact region-based representation that preserves object boundaries and local structures, and have therefore been widely used in a variety of vision tasks to reduce computational cost. However, most existing superpixel algorithms produce irregularly shaped regions, which are not well aligned with regular operators such as convolutions. Consequently, superpixels are often treated as an offline preprocessing step, limiting parallel implementation and hindering end-to-end optimization within deep learning pipelines. Motivated by the adaptive representation and coverage property of granular-ball computing, we develop a square superpixel generation approach. Specifically, we approximate superpixels using multi-scale square blocks to avoid the computational and implementation difficulties induced by irregular shapes, enabling efficient parallel processing and learnable feature extraction. For each block, a purity score is computed based on pixel-intensity similarity, and high-quality blocks are selected accordingly. The resulting square superpixels can be readily integrated as graph nodes in graph neural networks (GNNs) or as tokens in Vision Transformers (ViTs), facilitating multi-scale information aggregation and structured visual representation. Experimental results on downstream tasks demonstrate consistent performance improvements, validating the effectiveness of the proposed method.

cs.CV

Bergman kernels and Poincar\'e series

We show that the Bergman kernel of a finite-volume quotient of a Hermitian manifold $\widetilde{X}$ with bounded geometry by a discrete group $\Gamma$ of its isometries is the same as the averaging over $\Gamma$ of the Bergman kernel on $\widetilde{X}$. We then use these results when $\widetilde{X}$ is a Hermitian symmetric space to show that a large class of relative Poincar\'e series does not vanish. This extends the results of Borthwick-Paul-Uribe and Barron (formerly Foth) to the case of general locally symmetric spaces of finite volume.

math.DG

HyperAlign: Hyperbolic Entailment Cones for Adaptive Text-to-Image Alignment Assessment

With the rapid development of text-to-image generation technology, accurately assessing the alignment between generated images and text prompts has become a critical challenge. Existing methods rely on Euclidean space metrics, neglecting the structured nature of semantic alignment, while lacking adaptive capabilities for different samples. To address these limitations, we propose HyperAlign, an adaptive text-to-image alignment assessment framework based on hyperbolic entailment geometry. First, we extract Euclidean features using CLIP and map them to hyperbolic space. Second, we design a dynamic-supervision entailment modeling mechanism that transforms discrete entailment logic into continuous geometric structure supervision. Finally, we propose an adaptive modulation regressor that utilizes hyperbolic geometric features to generate sample-level modulation parameters, adaptively calibrating Euclidean cosine similarity to predict the final score. HyperAlign achieves highly competitive performance on both single database evaluation and cross-database generalization tasks, fully validating the effectiveness of hyperbolic geometric modeling for image-text alignment assessment.

cs.CV

DPTrack:Directional Kernel-Guided Prompt Learning for Robust Nighttime Aerial Tracking

Existing nighttime aerial trackers based on prompt learning rely solely on spatial localization supervision, which fails to provide fine-grained cues that point to target features and inevitably produces vague prompts. This limitation impairs the tracker's ability to accurately focus on the object features and results in trackers still performing poorly. To address this issue, we propose DPTrack, a prompt-based aerial tracker designed for nighttime scenarios by encoding the given object's attribute features into the directional kernel enriched with fine-grained cues to generate precise prompts. Specifically, drawing inspiration from visual bionics, DPTrack first hierarchically captures the object's topological structure, leveraging topological attributes to enrich the feature representation. Subsequently, an encoder condenses these topology-aware features into the directional kernel, which serves as the core guidance signal that explicitly encapsulates the object's fine-grained attribute cues. Finally, a kernel-guided prompt module built on channel-category correspondence attributes propagates the kernel across the features of the search region to pinpoint the positions of target features and convert them into precise prompts, integrating spatial gating for robust nighttime tracking. Extensive evaluations on established benchmarks demonstrate DPTrack's superior performance. Our code will be available at https://github.com/zzq-vipsl/DPTrack.

cs.CV

The rare decays $h^0\rightarrow Z\gamma,V Z$ in the NB-LSSM

This study investigates the Higgs rare decays $h^0\rightarrow Z\gamma,V Z$ within the next to minimum B-L supersymmetric model(NB-LSSM), where $V$ represents a vector meson $(\phi,J/\psi,\Upsilon(1S),\rho^0,\omega)$. Compared to the minimal supersymmetric standard model(MSSM), the NB-LSSM introduces three singlet Higgs superfields, which mix with the Higgs doublets and affect the lightest Higgs boson mass and the Higgs couplings. The loop-induced contributions resulting from the effective $h^0Z\gamma$ coupling can produce new physics(NP) contributions, thereby affecting the theoretical predictions of rare decays significantly through the new parameters such as $\kappa$, $\lambda$, $\lambda_2\cdot\cdot\cdot$. The results of this work can provide a reference for probing NP beyond standard model(SM).

hep-ph

FCA2: Frame Compression-Aware Autoencoder for Modular and Fast Compressed Video Super-Resolution

State-of-the-art (SOTA) compressed video super-resolution (CVSR) models face persistent challenges, including prolonged inference time, complex training pipelines, and reliance on auxiliary information. As video frame rates continue to increase, the diminishing inter-frame differences further expose the limitations of traditional frame-to-frame information exploitation methods, which are inadequate for addressing current video super-resolution (VSR) demands. To overcome these challenges, we propose an efficient and scalable solution inspired by the structural and statistical similarities between hyperspectral images (HSI) and video data. Our approach introduces a compression-driven dimensionality reduction strategy that reduces computational complexity, accelerates inference, and enhances the extraction of temporal information across frames. The proposed modular architecture is designed for seamless integration with existing VSR frameworks, ensuring strong adaptability and transferability across diverse applications. Experimental results demonstrate that our method achieves performance on par with, or surpassing, the current SOTA models, while significantly reducing inference time. By addressing key bottlenecks in CVSR, our work offers a practical and efficient pathway for advancing VSR technology. Our code will be publicly available at https://github.com/handsomewzy/FCA2.

eess.IV

EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment

Free-energy-guided self-repair mechanisms have shown promising results in image quality assessment (IQA), but remain under-explored in video quality assessment (VQA), where temporal dynamics and model constraints pose unique challenges. Unlike static images, video content exhibits richer spatiotemporal complexity, making perceptual restoration more difficult. Moreover, VQA systems often rely on pre-trained backbones, which limits the direct integration of enhancement modules without affecting model stability. To address these issues, we propose EyeSimVQA, a novel VQA framework that incorporates free-energy-based self-repair. It adopts a dual-branch architecture, with an aesthetic branch for global perceptual evaluation and a technical branch for fine-grained structural and semantic analysis. Each branch integrates specialized enhancement modules tailored to distinct visual inputs-resized full-frame images and patch-based fragments-to simulate adaptive repair behaviors. We also explore a principled strategy for incorporating high-level visual features without disrupting the original backbone. In addition, we design a biologically inspired prediction head that models sweeping gaze dynamics to better fuse global and local representations for quality prediction. Experiments on five public VQA benchmarks demonstrate that EyeSimVQA achieves competitive or superior performance compared to state-of-the-art methods, while offering improved interpretability through its biologically grounded design.

cs.CV

The Athenian Academy: A Seven-Layer Architecture Model for Multi-Agent Systems

This paper proposes the "Academy of Athens" multi-agent seven-layer framework, aimed at systematically addressing challenges in multi-agent systems (MAS) within artificial intelligence (AI) art creation, such as collaboration efficiency, role allocation, environmental adaptation, and task parallelism. The framework divides MAS into seven layers: multi-agent collaboration, single-agent multi-role playing, single-agent multi-scene traversal, single-agent multi-capability incarnation, different single agents using the same large model to achieve the same target agent, single-agent using different large models to achieve the same target agent, and multi-agent synthesis of the same target agent. Through experimental validation in art creation, the framework demonstrates its unique advantages in task collaboration, cross-scene adaptation, and model fusion. This paper further discusses current challenges such as collaboration mechanism optimization, model stability, and system security, proposing future exploration through technologies like meta-learning and federated learning. The framework provides a structured methodology for multi-agent collaboration in AI art creation and promotes innovative applications in the art field.

cs.MA

CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information

Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the valuable ``beyond-image-modality" information embedded in EEG signals. This results in the loss of critical multimodal information in EEG. To address this limitation, we propose CognitionCapturer, a unified framework that fully leverages multimodal data to represent EEG signals. Specifically, CognitionCapturer trains Modality Expert Encoders for each modality to extract cross-modal information from the EEG modality. Then, it introduces a diffusion prior to map the EEG embedding space to the CLIP embedding space, followed by using a pretrained generative model, the proposed framework can reconstruct visual stimuli with high semantic and structural fidelity. Notably, the framework does not require any fine-tuning of the generative models and can be extended to incorporate more modalities. Through extensive experiments, we demonstrate that CognitionCapturer outperforms state-of-the-art methods both qualitatively and quantitatively. Code: https://github.com/XiaoZhangYES/CognitionCapturer.

cs.CV

Stability equivalence for stochastic differential equations, stochastic differential delay equations and their corresponding Euler-Maruyama methods in $G$-framework

In this paper, we investigate the stability equivalence problem for stochastic differential delay equations, the auxiliary stochastic differential equations and their corresponding Euler-Maruyama (EM) methods under $G$-framework. More precisely, for $p\geq 2$, we prove the equivalence of practical exponential stability in $p$-th moment sense among stochastic differential delay equations driven by $G$-Brownian motion ($G$-SDDEs), the auxiliary stochastic differential equations driven by $G$-Brownian motion ($G$-SDEs), and their corresponding Euler-Maruyama methods, provided the delay or the step size is small enough. Thus, we can carry out careful simulations to examine the practical exponential stability of the underlying $G$-SDDE or $G$-SDE under some reasonable assumptions.

math.PR

Self-organized intracellular twisters

Life in complex systems, such as cities and organisms, comes to a standstill when global coordination of mass, energy, and information flows is disrupted. Global coordination is no less important in single cells, especially in large oocytes and newly formed embryos, which commonly use fast fluid flows for dynamic reorganization of their cytoplasm. Here, we combine theory, computing, and imaging to investigate such flows in the Drosophila oocyte, where streaming has been proposed to spontaneously arise from hydrodynamic interactions among cortically anchored microtubules loaded with cargo-carrying molecular motors. We use a fast, accurate, and scalable numerical approach to investigate fluid-structure interactions of 1000s of flexible fibers and demonstrate the robust emergence and evolution of cell-spanning vortices, or twisters. Dominated by a rigid body rotation and secondary toroidal components, these flows are likely involved in rapid mixing and transport of ooplasmic components.

physics.bio-ph

MKANet: A Lightweight Network with Sobel Boundary Loss for Efficient Land-cover Classification of Satellite Remote Sensing Imagery

Land cover classification is a multi-class segmentation task to classify each pixel into a certain natural or man-made category of the earth surface, such as water, soil, natural vegetation, crops, and human infrastructure. Limited by hardware computational resources and memory capacity, most existing studies preprocessed original remote sensing images by down sampling or cropping them into small patches less than 512*512 pixels before sending them to a deep neural network. However, down sampling images incurs spatial detail loss, renders small segments hard to discriminate, and reverses the spatial resolution progress obtained by decades of years of efforts. Cropping images into small patches causes a loss of long-range context information, and restoring the predicted results to their original size brings extra latency. In response to the above weaknesses, we present an efficient lightweight semantic segmentation network termed MKANet. Aimed at the characteristics of top view high-resolution remote sensing imagery, MKANet utilizes sharing kernels to simultaneously and equally handle ground segments of inconsistent scales, and also employs parallel and shallow architecture to boost inference speed and friendly support image patches more than 10X larger. To enhance boundary and small segments discrimination, we also propose a method that captures category impurity areas, exploits boundary information and exerts an extra penalty on boundaries and small segment misjudgment. Both visual interpretations and quantitative metrics of extensive experiments demonstrate that MKANet acquires state-of-the-art accuracy on two land-cover classification datasets and infers 2X faster than other competitive lightweight networks. All these merits highlight the potential of MKANet in practical applications.

cs.CV

Do Deep Neural Networks Always Perform Better When Eating More Data?

Data has now become a shortcoming of deep learning. Researchers in their own fields share the thinking that "deep neural networks might not always perform better when they eat more data," which still lacks experimental validation and a convincing guiding theory. Here to fill this lack, we design experiments from Identically Independent Distribution(IID) and Out of Distribution(OOD), which give powerful answers. For the purpose of guidance, based on the discussion of results, two theories are proposed: under IID condition, the amount of information determines the effectivity of each sample, the contribution of samples and difference between classes determine the amount of sample information and the amount of class information; under OOD condition, the cross-domain degree of samples determine the contributions, and the bias-fitting caused by irrelevant elements is a significant factor of cross-domain. The above theories provide guidance from the perspective of data, which can promote a wide range of practical applications of artificial intelligence.

cs.LG

Regularized Deep Linear Discriminant Analysis

As a non-linear extension of the classic Linear Discriminant Analysis(LDA), Deep Linear Discriminant Analysis(DLDA) replaces the original Categorical Cross Entropy(CCE) loss function with eigenvalue-based loss function to make a deep neural network(DNN) able to learn linearly separable hidden representations. In this paper, we first point out DLDA focuses on training the cooperative discriminative ability of all the dimensions in the latent subspace, while put less emphasis on training the separable capacity of single dimension. To improve DLDA, a regularization method on within-class scatter matrix is proposed to strengthen the discriminative ability of each dimension, and also keep them complement each other. Experiment results on STL-10, CIFAR-10 and Pediatric Pneumonic Chest X-ray Dataset showed that our proposed regularization method Regularized Deep Linear Discriminant Analysis(RDLDA) outperformed DLDA and conventional neural network with CCE as objective. To further improve the discriminative ability of RDLDA in the local space, an algorithm named Subclass RDLDA is also proposed.

cs.CV

Bergman kernels and equidistribution for sequences of line bundles on Kähler manifolds

Given a sequence of positive Hermitian holomorphic line bundles $(L_p,h_p)$ on a Kähler manifold $X$, we establish the asymptotic expansion of the Bergman kernel of the space of global holomorphic sections of $L_p$, under a natural convergence assumption on the sequence of curvatures $c_1(L_p,h_p)$. We then apply this to study the asymptotic distribution of common zeros of random sequences of $m$-tuples of sections of $L_p$ as $p\to\infty$.

math.CV