SearcharxivSearch

arXiv subjects

Jie Wen

Publications and source records attributed to Jie Wen.

At least 19 recordsLinked to original sources

When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning

Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observed views, but also label-aware exploitation of view-specific evidence. Existing approaches usually emphasize either shared representation learning or decision-level fusion. The former improves robustness against missing views, yet tends to compress label-discriminative view-specific cues into a single latent representation. The latter preserves individual view predictions, but often relies on fixed or globally learned fusion weights, ignoring that different labels of different instances may require different views. To address these limitations, this paper presents V2L, a unified representation-decision framework for incomplete multi-view multi-label classification. On the representation side, V2L constructs semantically consistent variational posteriors from incomplete views through a perturbation-aware encoding mechanism, which provides a stable shared semantic basis. On the decision side, V2L introduces an active view-label relevance modeling strategy that estimates instance-wise and label-wise view contributions, allowing each label prediction to adaptively select useful view-specific evidence. From the perspective of model architecture, these two important strategies are integrated into a unified framework through a hybrid fusion architecture, simultaneously meeting the requirements of cross-view semantic consistency and representational complementarity. Extensive experiments under both incomplete and complete settings show that V2L achieves leading performance on five benchmarks. Code is available at: https://github.com/justsmart/V2L.

cs.CV

On the Frankl--Tokushige conjecture and almost complete $r$-cross $t$-intersection theorems for vector spaces

Let $r\geq3$ and $k_1\geq k_2\geq\cdots\geq k_r\geq t$. Let $\mathcal{F}_1,\mathcal{F}_2,\ldots,\mathcal{F}_r$ be families of subspaces, of respective dimensions $k_1,k_2,\ldots,k_r$, in an $n$-dimensional vector space over the finite field $\mathbb{F}_q$. The $r$ families are called $r$-cross $t$-intersecting if $\dim \left(F_{1} \cap F_{2} \cap \cdots \cap F_{r}\right) \geq t$ for all $F_{i} \in \mathcal{F}_{i}, i = 1,2,\dots,r$. In 2016, Frankl and Tokushige conjectured that $\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-1\brack k_i-1}$ for $t=1$ and $n\geq rk_1/(r-1)$. The appealing conjecture suggests establishing intersection theorems for $n\sim ck_1$ with $c=c(r)\in(1,2)$, a direction that has long been challenging. In this paper, we overcome this barrier by proving that $$\prod_{i=1}^{r}|\mathcal{F}_i|\leq\prod_{i=1}^{r}{n-t\brack k_i-t}\;\;\mbox{for all}\;\;t\geq1\;\mbox{and}\;n\geq rk_1/(r-1)+C(t,r),$$ where $C(t,r)=rt/(r-1)+1$. This proves the Frankl--Tokushige conjecture except for at most three values of $n$, and establishes an Erd\H{o}s--Ko--Rado type theorem for almost all values of parameters. Furthermore, we characterize all extremal configurations. Our proof is purely combinatorial and based on the $t$-cover method, with several essential refinements. We also obtain almost complete intersection theorems for $r$-wise $t$-intersecting families and non-trivial $r$-cross $t$-intersecting families.

math.CO

Structure of large $t$-intersecting families I: Stability for the Hilton--Milner--Frankl theorem

We study the structure of large $t$-intersecting families. A family of $k$-subsets of an $n$-set is $t$-intersecting if every two of its members intersect in at least $t$ elements. A $t$-intersecting family is non-trivial if no $t$-subset is contained in all its members. We prove several stability results for the seminal Hilton--Milner--Frankl theorem. First, for any fixed $\eta,\varepsilon,\theta\in(0,1)$, we prove that if $k/t\geq1+\eta$ and $n=\Omega(tk^{1+\varepsilon})$, then every non-trivial $t$-intersecting family of size greater than $(1+\theta)|\mathcal{K}|$ is a subfamily of one of the two extremal families in the theorem, where $\mathcal{K}$ is an explicit large non-trivial $t$-intersecting family. The key ingredient in the proof is a removal lemma. We also obtain a classification of all $t$-intersecting families with size bounded below by $|\mathcal{K}|$ minus an explicit lower-order term, provided that $k\geq t+4\geq6$ and $n\geq t+6\cdot\max\{(t+2)^2, k(k-t)\}$. This strengthens results of Cao--Lv--Wang (2021) and Frankl (2025) for a broad range of $k$ and $t$ (for example, when $k-t\geq2\sqrt{t}$). As an application of this classification, we determine the largest $t$-intersecting families for each prescribed lower bound on $t$-diversity not exceeding $t(n-k)$, thereby obtaining $t$-intersection versions of results of Han and Kohayakawa (2017) and Kupavskii (2025). To establish these results, we develop techniques based on the spread approximation method and the $t$-cover method, which may be useful for other intersection problems.

math.CO

A unified approach to cross-intersection problems with applications to Hilton--Milner type theorems and stability

We develop a new approach to cross-intersection problems in extremal set theory. The method builds on the iterative procedure introduced by Kupavskii and Zakharov (2024) and the $t$-cover method. It provides a flexible framework for deriving extremal and stability results for cross $t$-intersecting families. Our approach applies to a variety of combinatorial objects. As an application, we prove a product version of the seminal Erd\H{o}s--Ko--Rado theorem for sufficiently spread set systems. Two families $\mathcal{F}$ and $\mathcal{G}$ of $k$-subsets of $[n]$ are called cross $t$-intersecting if $|F\cap G|\geq t$ for all $F\in\mathcal{F}$ and $G\in\mathcal{G}$. We determine the families maximizing $\min\{|\mathcal{F}|, |\mathcal{G}|\}$ for large $n$ and all $t\ge2$, generalizing results of M\"{o}rs (1985) and F\"{u}redi (1995) for cross $1$-intersecting families. We then determine the families maximizing $|\mathcal{F}||\mathcal{G}|$ under the condition $\max\{|\cap_{F\in\mathcal{F}}F|,|\cap_{G\in\mathcal{G}}G|\}<t$ for large $n$. This improves the bound obtained by Frankl and Wang (2024), and provides a characterization of extremal configurations. For a family $\mathcal{F}$ of subsets of $[n]$, we introduce its $t$-diversity $\gamma_t(\mathcal{F})$, defined as the minimum number of sets from $\mathcal{F}$ not containing a fixed $t$-subset. This serves as a natural generalization of the important notion of diversity for $t=1$. We obtain a stability result via $\gamma_t$, and determine the maximum of $\min\{\gamma_t(\mathcal{F}),\gamma_t(\mathcal{G})\}$ for cross $t$-intersecting families $\mathcal{F}$ and $\mathcal{G}$. These yield new results for $t$-intersecting families, including a stability theorem towards a conjecture of Ellis, Keller and Lifshitz (2019), which may also be regarded as a $t$-intersection version, for large $n$, of an influential theorem of Frankl (1987).

math.CO

Expandable, Compressible, Mineable: Open-World Thermal Image Restoration

In open-world settings, thermal infrared (TIR) image degradations continuously emerge and evolve, while most existing all-in-one restoration methods are built on a closed-set assumption and struggle to continually adapt to novel degradations. To address this, we propose ECMRNet, an Expandable, Compressible, and Mineable Restoration Network for open-world TIR restoration from a continual learning perspective. Conceptually, ECMRNet unifies continual degradation learning as an "expand-compress-mine" closed-loop process, enabling sustained adaptation to new degradations with controllable evolution. Structurally, ECMRNet decomposes intermediate representations into group-isolated subspaces, and achieves strict parameter isolation and fast adaptation to new degradations by freezing historical groups and isomorphically expanding new ones. To curb model growth as tasks accumulate, we present Structural Entropy Pruning, which identifies and removes redundant channel groups via two-dimensional structural entropy minimization, achieving information contribution-driven adaptive compression. Moreover, we design a Sub-degradation Knowledge Mining Module that dynamically retrieves and recombines transferable components from historical representations to improve restoration under compound degradations. Experimental results demonstrate that ECMRNet achieves superior overall performance across diverse single and compound degradations while using fewer parameters and lower computational cost. The source code is available at https://github.com/Kust-lp/ECMRNet.

cs.CV

SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere

Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods universally depend on large-scale annotations or task-specific training procedures, severely limiting their rapid deployment to novel scenes. We observe that intermediate-layer features of pre-trained multimodal large language models (MLLMs) already encode rich anomaly semantics, yet existing approaches rely on the language output pathway and fail to exploit the geometric discriminability latent in these representations. Based on this finding, we propose SphereVAD, a fully training-free, zero-shot VAD framework that recasts anomaly discrimination as von Mises-Fisher (vMF) likelihood-ratio geodesic inference on the unit hypersphere, unleashing latent discriminability through principled geometric reasoning rather than learning new representations. Specifically, SphereVAD first applies Frechet mean centering to unfold feature distributions and eliminate domain biases, then employs Holistic Scene Attention (HSA) to reinforce feature consistency using cross-video priors, and finally performs vMF-guided Spherical Geodesic Pulling (SGP) to align ambiguous segments with directional prototypes on the spherical manifold. This training-free pipeline requires only minimal synthetic images for calibration. SphereVAD establishes new state-of-the-art results among training-free approaches on three major benchmarks and remains competitive with fully supervised baselines. Code will be available upon acceptance.

cs.CV

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills

Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering into cross-file security review. Existing guardrails often flag risk but recover malicious intent inconsistently under semantics-preserving rewrites. This paper formulates pre-load auditing for untrusted Agent Skills as a robust three-way classification task and introduces SkillGuard-Robust, which combines role-aware evidence extraction, selective semantic verification, and consistency-preserving adjudication. We evaluate SkillGuard-Robust on SkillGuardBench and two public-ecosystem extensions through five large evaluation views ranging from 254 to 404 packages. On the 404-package held-out aggregate, SkillGuard-Robust reaches 97.30% overall exact match, 98.33% malicious-risk recall, and 98.89% attack exact consistency. On the 254-package external-ecosystem view, it reaches 99.66%, 100.00%, and 100.00%, respectively. These results support a bounded conclusion: factorized package auditing materially improves frozen and public-ecosystem robustness, while harsher external-source transfer remains an open challenge.

cs.CR

Feature-Label Modal Alignment for Robust Partial Multi-Label Learning

In partial multi-label learning (PML), each instance is associated with a set of candidate labels containing both ground-truth and noisy labels. The presence of noisy labels disrupts the correspondence between features and labels, degrading classification performance. To address this challenge, we propose a novel PML method based on feature-label modal alignment (PML-MA), which treats features and labels as two complementary modalities and restores their consistency through systematic alignment. Specifically, PML-MA first employs low-rank orthogonal decomposition to generate pseudo-labels that approximate the true label distribution by filtering noisy labels. It then aligns features and pseudo-labels through both global projection into a common subspace and local preservation of neighborhood structures. Finally, a multi-peak class prototype learning mechanism leverages the multi-label nature where instances simultaneously belong to multiple categories, using pseudo-labels as soft membership weights to enhance discriminability. By integrating modal alignment with prototype-guided refinement, PML-MA ensures pseudo-labels better reflect the true distribution while maintaining robustness against label noise. Extensive experiments on both real-world and synthetic datasets demonstrate that PML-MA significantly outperforms state-of-the-art methods, achieving superior classification accuracy and noise robustness.

cs.LG

SEHFS: Structural Entropy-Guided High-Order Correlation Learning for Multi-View Multi-Label Feature Selection

In recent years, multi-view multi-label learning (MVML) has attracted extensive attention due to its close alignment to real-world scenarios. Information-theoretic methods have gained prominence for learning nonlinear correlations. However, two key challenges persist: first, features in real-world data commonly exhibit high-order structural correlations, but existing information-theoretic methods struggle to learn such correlations; second, commonly relying on heuristic optimization, information-theoretic methods are prone to converging to local optima. To address these two challenges, we propose a novel method called Structural Entropy Guided High-Order Correlation Learning for Multi-View Multi-Label Feature Selection (SEHFS). The core idea of SEHFS is to convert the feature graph into a structural-entropy-minimizing encoding tree, quantifying the information cost of high-order dependencies and thus learning high-order feature correlations beyond pairwise correlations. Specifically, features exhibiting strong high-order redundancy are grouped into a single cluster within the encoding tree, while inter-cluster feaeture correlations are minimized, thereby eliminating redundancy both within and across clusters. Furthermore, a new framework based on the fusion of information theory and matrix methods is adopted, which learns a shared semantic matrix and view-specific contribution matrices to reconstruct a global view matrix, thereby enhancing the information-theoretic method and balancing the global and local optimization. The ability of structural entropy to learn high-order correlations is theoretically established, and and both experiments on eight datasets from various domains and ablation studies demonstrate that SEHFS achieves superior performance in feature selection.

cs.LG

On $r$-cross $t$-intersecting families of partitions

In this paper, we address several intersection problems for $r$-cross $t$-intersecting families of partitions. A $k$-partition of an $n$-set $X$ is a set of $k$ pairwise disjoint non-empty subsets whose union is $X$. For $1\leq i\leq r$, let $\mathcal{F}_i$ be a family of $k_i$-partitions of $X$. We say that $\mathcal{F}_1,\mathcal{F}_2,\ldots,\mathcal{F}_r$ are $r$-cross $t$-intersecting if $|\cap_{i=1}^{r}F_i|\geq t$ for all $F_i\in\mathcal{F}_i$. The families are called non-trivial if $|\cap_{i=1}^r(\cap_{F\in\mathcal{F}_i}F)|<t$. Proving an Erd\H{o}s-Ko-Rado type theorem, we determine the families maximizing $\prod_{i=1}^r|\mathcal{F}_i|$. We further determine non-trivial $r$-cross $t$-intersecting families with maximum product of sizes; this result also serves as a Hilton-Milner type theorem. In particular, for $r=2$ there are two potential structures for optimal families, and for $r\geq3$ exactly one remains.

math.CO

Prompt Tuning for CLIP on the Pretrained Manifold

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold toward directions that are unfavorable for transfer. This drift degrades generalization. To address this limitation, we propose ManiPT, a framework that performs prompt tuning on the pretrained manifold. ManiPT introduces cosine consistency constraints in both the text and image modalities to confine the learned representations within the pretrained geometric neighborhood. Furthermore, we introduce a structural bias that enforces incremental corrections, guiding the adaptation along transferable directions to mitigate reliance on shortcut learning. From a theoretical perspective, ManiPT alleviates overfitting tendencies under limited data. Our experiments cover four downstream settings: unseen-class generalization, few-shot classification, cross-dataset transfer, and domain generalization. Across these settings, ManiPT achieves higher average performance than baseline methods. Notably, ManiPT provides an explicit perspective on how prompt tuning overfits under limited supervision.

cs.CV

M3-AD: Reflection-aware Multi-modal, Multi-category, and Multi-dimensional Benchmark and Framework for Industrial Anomaly Detection

Although multimodal large language models (MLLMs) have advanced industrial anomaly detection toward a zero-shot paradigm, they still tend to produce high-confidence yet unreliable decisions in fine-grained and structurally complex industrial scenarios, and lack effective self-corrective mechanisms. To address this issue, we propose M3-AD, a unified reflection-aware multimodal framework for industrial anomaly detection. M3-AD comprises two complementary data resources: M3-AD-FT, designed for reflection-aligned fine-tuning, and M3-AD-Bench, designed for systematic cross-category evaluation, together providing a foundation for reflection-aware learning and reliability assessment. Building upon this foundation, we propose RA-Monitor, which models reflection as a learnable decision revision process and guides models to perform controlled self-correction when initial judgments are unreliable, thereby improving decision robustness. Extensive experiments conducted on M3-AD-Bench demonstrate that RA-Monitor outperforms multiple open-source and commercial MLLMs in zero-shot anomaly detection and anomaly analysis tasks. Code will be released at https://github.com/Yanhui-Lee/M3-AD.

cs.LG

Q-Hawkeye: Reliable Visual Policy Optimization for Image Quality Assessment

Image Quality Assessment (IQA) predicts perceptual quality scores consistent with human judgments. Recent RL-based IQA methods built on MLLMs focus on generating visual quality descriptions and scores, ignoring two key reliability limitations: (i) although the model's prediction stability varies significantly across training samples, existing GRPO-based methods apply uniform advantage weighting, thereby amplifying noisy signals from unstable samples in gradient updates; (ii) most works emphasize text-grounded reasoning over images while overlooking the model's visual perception ability of image content. In this paper, we propose Q-Hawkeye, an RL-based reliable visual policy optimization framework that redesigns the learning signal through unified Uncertainty-Aware Dynamic Optimization and Perception-Aware Optimization. Q-Hawkeye estimates predictive uncertainty using the variance of predicted scores across multiple rollouts and leverages this uncertainty to reweight each sample's update strength, stabilizing policy optimization. To strengthen perceptual reliability, we construct paired inputs of degraded images and their original images and introduce an Implicit Perception Loss that constrains the model to ground its quality judgments in genuine visual evidence. Extensive experiments demonstrate that Q-Hawkeye outperforms state-of-the-art methods and generalizes better across multiple datasets. Our dataset and code are available at https://github.com/AMAP-ML/Q-Hawkeye.

cs.CV

Advancing Adaptive Multi-Stage Video Anomaly Reasoning: A Benchmark Dataset and Method

Recent progress in reasoning capabilities of Multimodal Large Language Models(MLLMs) has highlighted their potential for performing complex video understanding tasks. However, in the domain of Video Anomaly Detection and Understanding (VAD&U), existing MLLM-based methods are largely limited to anomaly localization or post-hoc description, lacking explicit reasoning processes, risk awareness, and decision-oriented interpretation. To address this gap, we define a new task termed Video Anomaly Reasoning (VAR), which elevates video anomaly analysis from descriptive understanding to structured, multi-stage reasoning. VAR explicitly requires models to perform progressive reasoning over anomalous events before answering anomaly-related questions, encompassing visual perception, causal interpretation, and risk-aware decision making. To support this task, we present a new dataset with 8,641 videos, where each video is annotated with diverse question types corresponding to different reasoning depths, totaling more than 50,000 samples, making it one of the largest datasets for video anomaly. The annotations are based on a structured Perception-Cognition-Action Chain-of-Thought (PerCoAct-CoT), which formalizes domain-specific reasoning priors for video anomaly understanding. This design enables systematic evaluation of multi-stage and adaptive anomaly reasoning. In addition, we propose Anomaly-Aware Group Relative Policy Optimization to further enhance reasoning reliability under weak supervision. Building upon the proposed task and dataset, we develop an end-to-end MLLM-based VAR model termed Vad-R1-Plus, which supports adaptive hierarchical reasoning and risk-aware decision making. Extensive experiments demonstrate that the proposed benchmark and method effectively advance the reasoning capabilities of MLLMs on VAR tasks, outperforming both open-source and proprietary baselines.

cs.CV

Diffusion Model Regularized Implicit Neural Representation for CT Metal Artifact Reduction

Computed tomography (CT) images are often severely corrupted by artifacts in the presence of metals. Existing supervised metal artifact reduction (MAR) approaches suffer from performance instability on known data due to their reliance on limited paired metal-clean data, which limits their clinical applicability. Moreover, existing unsupervised methods face two main challenges: 1) the CT physical geometry is not effectively incorporated into the MAR process to ensure data fidelity; 2) traditional heuristics regularization terms cannot fully capture the abundant prior knowledge available. To overcome these shortcomings, we propose diffusion model regularized implicit neural representation framework for MAR. The implicit neural representation integrates physical constraints and imposes data fidelity, while the pre-trained diffusion model provides prior knowledge to regularize the solution. Experimental results on both simulated and clinical data demonstrate the effectiveness and generalization ability of our method, highlighting its potential to be applied to clinical settings.

cs.CV

Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval

Domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, enabling effective retrieval while mitigating domain discrepancies. However, existing methods encounter several fundamental limitations: 1) neglecting class-level semantic alignment and excessively pursuing pair-wise sample alignment; 2) lacking either pseudo-label reliability consideration or geometric guidance for assessing label correctness; 3) directly quantizing original features affected by domain shift, undermining the quality of learned hash codes. In view of these limitations, we propose Prototype-Based Semantic Consistency Alignment (PSCA), a two-stage framework for effective domain adaptive retrieval. In the first stage, a set of orthogonal prototypes directly establishes class-level semantic connections, maximizing inter-class separability while gathering intra-class samples. During the prototype learning, geometric proximity provides a reliability indicator for semantic consistency alignment through adaptive weighting of pseudo-label confidences. The resulting membership matrix and prototypes facilitate feature reconstruction, ensuring quantization on reconstructed rather than original features, thereby improving subsequent hash coding quality and seamlessly connecting both stages. In the second stage, domain-specific quantization functions process the reconstructed features under mutual approximation constraints, generating unified binary hash codes across domains. Extensive experiments validate PSCA's superior performance across multiple datasets.

cs.LG

Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-Training

Multimodal protein features play a crucial role in protein function prediction. However, these features encompass a wide range of information, ranging from structural data and sequence features to protein attributes and interaction networks, making it challenging to decipher their complex interconnections. In this work, we propose a multimodal protein function prediction method (DSRPGO) by utilizing dynamic selection and reconstructive pre-training mechanisms. To acquire complex protein information, we introduce reconstructive pre-training to mine more fine-grained information with low semantic levels. Moreover, we put forward the Bidirectional Interaction Module (BInM) to facilitate interactive learning among multimodal features. Additionally, to address the difficulty of hierarchical multi-label classification in this task, a Dynamic Selection Module (DSM) is designed to select the feature representation that is most conducive to current protein function prediction. Our proposed DSRPGO model improves significantly in BPO, MFO, and CCO on human datasets, thereby outperforming other benchmark models.

cs.LG

Erdős-Ko-Rado theorem and Hilton-Milner type theorem for $k$-partitions

A $k$-partition of an $n$-set $X$ is a collection of $k$ pairwise disjoint non-empty subsets whose union is $X$. A family of $k$-partitions of $X$ is called $t$-intersecting if any two of its members share at least $t$ blocks. A $t$-intersecting family is trivial if every $k$-partition in it contains $t$ fixed blocks, and is non-trivial otherwise. In this paper, we first prove that, for $n\geq L(k,t):=(t+1)+(k-t+1)\cdot\log_2(t+1)(k-t+1)$, a $t$-intersecting family with maximum size must consist of all $k$-partitions containing $t$ fixed singletons. This improves the results given by Erdős and Székely (2000), and by Kupavskii (2023). We further determine the non-trivial $t$-intersecting families of $k$-partitions with maximum size for $n \ge 2L(k,t)$, which turn out to be natural analogs of the corresponding families for finite sets. In addition, we prove a stability result.

math.CO