SearcharxivSearch

arXiv subjects

Xiangyu Zhou

Publications and source records attributed to Xiangyu Zhou.

At least 19 recordsLinked to original sources

A logarithmic Bogomolov--Sommese vanishing theorem on compact K\"ahler manifolds

In this paper, we establish a logarithmic Bogomolov--Sommese vanishing theorem in terms of numerical dimension for pseudo-effective line bundles on compact K\"ahler manifolds. As an application, we obtain a rigidity result with vanishing second Chern class for logarithmic cotangent bundles by combining the vanishing theorem with a structure theorem of Iwai and Matsumura.

math.AG

On the extension of K\"ahler currents on compact complex manifolds

Let $(X,\omega)$ be a compact K\"ahler manifold and let $V\subset X$ be a closed complex submanifold. Coman-Guedj-Zeriahi proposed the problem: is every $\omega|_V$-plurisubharmonic function on $V$ the restriction of an $\omega$-plurisubharmonic function on $X$? In this paper, we solve this problem affirmatively, even for a compact Hermitian manifold.

math.CV

A remark on Bogomolov type vanishing theorem

Several vanishing theorems for pseudo-effective line bundles are presented. All related results in the present note can be derived from already known theorems, in particular from the Steenbrink vanishing theorem and Boucksom's generalization of Bogomolov's vanishing theorem. No originality or priority is claimed for these statements; they are listed here merely as conjugate forms of the Kawamata--Viehweg vanishing theorem, for convenience of reference and comparison with existing literature.

math.AG

Capacity Stability of Complex Monge-Amp\`ere Equations with Moving Prescribed Singularities

For complex Monge-Amp\`ere equations with moving big cohomology classes and prescribed model singularities of positive Monge-Amp\`ere mass, we prove that, under total variation convergence of the right-hand side non-pluripolar positive Radon measures, convergence of the prescribed model potentials in Monge-Amp\`ere capacity is equivalent to convergence in capacity of the associated normalized solutions. We further prove that the ceiling operator coincides with the singularity envelope for potentials associated to a big $(1,1)$-class, regardless of their Monge-Amp\`ere mass, thereby resolving a conjecture of Darvas-Di Nezza-Lu. Consequently, the singularity envelope is idempotent without the positivity assumption on the mass.

math.CV

GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness

Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start. This paper asks whether the order in which parameters become optimizable can affect the final robust solution, even when the final architecture or computation budget is controlled. We propose GRAPE, Guided Parameter-Space Evolution, a training framework for compact adversarial robustness. GRAPE combines parameter-space stabilization with progressive hidden expansion: it stabilizes robust optimization in the currently exposed space, gradually releases new optimizable dimensions, and uses an adversarial spectral utilization score to guide newly released capacity toward high-pressure modules. In contrast to fixed-structure AT, GRAPE treats robust model learning as a process of progressive parameter-space exposure and evolution. Under the standard $\ell_\infty$ threat model on CIFAR-10, with fixed-structure ResNet-18 AT as a controlled reference, GRAPE improves PGD-20 robust accuracy from 51.70% to 56.94% at a nearly matched computation budget with a FLOPs ratio of 1.009x, while reducing parameter count by about 21.4%. A sequential grow variant with the same final ResNet-18 architecture reaches 56.52% PGD-20 robust accuracy, indicating that the gain is not only due to final architecture differences but also to the parameter-space exposure path. These results suggest that guided parameter-space evolution can yield compact and robust parameter configurations under matched computation.

cs.LG

Macaulay representation of the prolongation matrix and the SOS conjecture

Let $z \in \mathbb{C}^n$, and let $A(z,\bar{z})$ be a real valued diagonal bihomogeneous Hermitian polynomial such that $A(z,\bar{z})\|z\|^2$ is a sum of squares, where $\|z\|$ denotes the Euclidean norm of $z$. In this paper, we provide an estimate for the rank of the sum of squares $A(z,\bar{z})\|z\|^2$ when $A(z,\bar{z})$ is not semipositive definite. As a consequence, we confirm the SOS conjecture proposed by Ebenfelt for $4 \leq n \leq 6$ when $A(z,\bar{z})$ is a real valued diagonal (not necessarily bihomogeneous) Hermitian polynomial, and we also give partial answers to the SOS conjecture for $n\geq 7$.

math.CV

A Newton-Okounkov Body Viewpoint on the SOS Conjecture

Let $z\in \mathbb C^n$ be the complex coordinates on $\mathbb C^n$, and $A(z,\bar z)$ be a real-valued Hermitian polynomial. The famous Ebenfelt's SOS conjecture asks for the minimum rank of $A(z,\bar z)\|z\|^2$ under the restriction that $A(z,\bar z)\|z\|^2$ is an SOS. Assume that $A(z,\bar z)$ is bihomogeneous. In the present note, we establish a connection between Ebenfelt's (Weak) SOS Conjecture and the theory of Newton-Okounkov bodies. By reformulating the conjecture in terms of lattice semigroups and their associated Newton-Okounkov convex bodies, we transform the problem of finding the minimal rank of a prolonged sum-of-squares polynomial into an extremal problem in convex geometry. In particular, we prove that this minimal rank is attained at the extreme points of a specific Newton-Okounkov body. Furthermore, if $A(z,\bar z)$ is moreover diagonal, we demonstrate that the relevant extreme points are finitely many rational points, thereby reducing the verification of the conjecture to a computationally tractable problem. This work provides a new tool for attacking the SOS Conjecture.

math.CV

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can describe visual content, their lack of explicit grounding leads to object hallucinations and unreliable depth reasoning, limiting their usefulness for accessibility guidance. We introduce WalkGPT, a pixel-grounded LVLM for the new task of Grounded Navigation Guide, unifying language reasoning and segmentation within a single architecture for depth-aware accessibility guidance. Given a pedestrian-view image and a navigation query, WalkGPT generates a conversational response with segmentation masks that delineate accessible and harmful features, along with relative depth estimation. The model incorporates a Multi-Scale Query Projector (MSQP) that shapes the final image tokens by aggregating them along text tokens across spatial hierarchies, and a Calibrated Text Projector (CTP), guided by a proposed Region Alignment Loss, that maps language embeddings into segmentation-aware representations. These components enable fine-grained grounding and depth inference without user-provided cues or anchor points, allowing the model to generate complete and realistic navigation guidance. We also introduce PAVE, a large-scale benchmark of 41k pedestrian-view images paired with accessibility-aware questions and depth-grounded answers. Experiments show that WalkGPT achieves strong grounded reasoning and segmentation performance. The source code and dataset are available on the \href{https://sites.google.com/view/walkgpt-26/home}{project website}.

cs.CV

Attention Smoothing Is All You Need For Unlearning

Large Language Models are prone to memorizing sensitive, copyrighted, or hazardous content, posing significant privacy and legal concerns. Retraining from scratch is computationally infeasible, whereas current unlearning methods exhibit unstable trade-offs between forgetting and utility, frequently producing incoherent outputs on forget prompts and failing to generalize due to the persistence of lexical-level and semantic-level associations in attention. We propose Attention Smoothing Unlearning (ASU), a principled framework that casts unlearning as self-distillation from a forget-teacher derived from the model's own attention. By increasing the softmax temperature, ASU flattens attention distributions and directly suppresses the lexical-level and semantic-level associations responsible for reconstructing memorized knowledge. This results in a bounded optimization objective that erases factual information yet maintains coherence in responses to forget prompts. Empirical evaluation on TOFU, MUSE, and WMDP, along with real-world and continual unlearning scenarios across question answering and text completion, demonstrates that ASU outperforms the baselines for most unlearning scenarios, delivering robust unlearning with minimal loss of model utility.

cs.LG

Owen-based Semantics and Hierarchy-Aware Explanation (O-Shap)

Shapley value-based methods have become foundational in explainable artificial intelligence (XAI), offering theoretically grounded feature attributions through cooperative game theory. However, in practice, particularly in vision tasks, the assumption of feature independence breaks down, as features (i.e., pixels) often exhibit strong spatial and semantic dependencies. To address this, modern SHAP implementations now include the Owen value, a hierarchical generalization of the Shapley value that supports group attributions. While the Owen value preserves the foundations of Shapley values, its effectiveness critically depends on how feature groups are defined. We show that commonly used segmentations (e.g., axis-aligned or SLIC) violate key consistency properties, and propose a new segmentation approach that satisfies the $T$-property to ensure semantic alignment across hierarchy levels. This hierarchy enables computational pruning while improving attribution accuracy and interpretability. Experiments on image and tabular datasets demonstrate that O-Shap outperforms baseline SHAP variants in attribution precision, semantic coherence, and runtime efficiency, especially when structure matters.

cs.AI

Vanishing theorems for pseudo-effective line bundles

In the present paper, we establish a general Kawamata-Viehweg-Koll\'ar-Nadel type vanishing theorem for higher direct images in terms of numerical dimension for closed positive currents on compact K\"ahler manifolds, unifying a number of important vanishing theorems.

math.CV

Attention Retention for Continual Learning with Vision Transformers

Continual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the attention to previously learned visual concepts shifts significantly after learning new tasks. Inspired by neuroscientific insights into the selective attention in the human visual system, we propose a novel attention-retaining framework to mitigate forgetting in CL. Our method constrains attention drift by explicitly modifying gradients during backpropagation through a two-step process: 1) extracting attention maps of the previous task using a layer-wise rollout mechanism and generating instance-adaptive binary masks, and 2) when learning a new task, applying these masks to zero out gradients associated with previous attention regions, thereby preventing disruption of learned visual concepts. For compatibility with modern optimizers, the gradient masking process is further enhanced by scaling parameter updates proportionally to maintain their relative magnitudes. Experiments and visualizations demonstrate the effectiveness of our method in mitigating catastrophic forgetting and preserving visual concepts. It achieves state-of-the-art performance and exhibits robust generalizability across diverse CL scenarios.

cs.CV

Not All Tokens Are Meant to Be Forgotten

Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they tend to memorize unwanted information, such as private or copyrighted content, raising significant privacy and legal concerns. Unlearning has emerged as a promising solution, but existing methods face a significant challenge of over-forgetting. This issue arises because they indiscriminately suppress the generation of all the tokens in forget samples, leading to a substantial loss of model utility. To overcome this challenge, we introduce the Targeted Information Forgetting (TIF) framework, which consists of (1) a flexible targeted information identifier designed to differentiate between unwanted words (UW) and general words (GW) in the forget samples, and (2) a novel Targeted Preference Optimization approach that leverages Logit Preference Loss to unlearn unwanted information associated with UW and Preservation Loss to retain general information in GW, effectively improving the unlearning process while mitigating utility degradation. Extensive experiments on the TOFU and MUSE benchmarks demonstrate that the proposed TIF framework enhances unlearning effectiveness while preserving model utility and achieving state-of-the-art results.

cs.LG

Degenerate Complex Hessian type equations on compact Hermitian manifolds and Applications

The aim of this paper is to further develop the theory of the degenerate complex Hessian equations on compact Hermitian manifolds. Building upon the generalization of the Bedford-Taylor pluripotential theory to complex Hessian equations by Kołodziej-Nguyen, we solve these equations in the $(ω, m)$-positive cone, $(ω, m)$-big classes and in nef classes, where $ω$ is a reference Hermitian metric. These results are also new in the Kähler case. Moreover, we adapt our techniques to solve complex Monge-Ampère equations in nef classes with mild singularities. The solutions we obtain, in the compact Kähler case, coincide with those for the complex Monge-Ampère equations in the sense of the non-pluripolar product introduced by Boucksom-Eyssidieux-Guedj-Zeriahi. One of the key ingredients in the proof is the adaption, to the Hermitian setting, of a new a priori $L^\infty$-estimate established by Guo-Phong-Tong and Guo-Phong-Tong-Wang.

math.CV

LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consistency. Existing T2I models often struggle with tasks like accurate text rendering, precise pose generation, or intricate compositional coherence. Concurrently, Vision-Language Models (LVLMs) have demonstrated powerful capabilities in cross-modal understanding and instruction following. We propose LumiGen, a novel LVLM-enhanced iterative framework designed to elevate T2I model performance, particularly in areas requiring fine-grained control, through a closed-loop, LVLM-driven feedback mechanism. LumiGen comprises an Intelligent Prompt Parsing & Augmentation (IPPA) module for proactive prompt enhancement and an Iterative Visual Feedback & Refinement (IVFR) module, which acts as a "visual critic" to iteratively correct and optimize generated images. Evaluated on the challenging LongBench-T2I Benchmark, LumiGen achieves a superior average score of 3.08, outperforming state-of-the-art baselines. Notably, our framework demonstrates significant improvements in critical dimensions such as text rendering and pose expression, validating the effectiveness of LVLM integration for more controllable and higher-quality image generation.

cs.LG

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning

Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process. Existing CoT compression methods mainly rely on global length rewards, which conflate necessary intermediate reasoning with redundant text and may therefore compromise reasoning fidelity. This paper revisits overthinking from a semantic-efficiency perspective and decomposes CoT redundancy into two distinct forms: internal redundancy, defined as informational stagnation before the first correct answer, and external redundancy, defined as superfluous continuation after the first correct answer. Based on this decomposition, we propose a dual-penalty reinforcement learning framework that separately optimizes reasoning progress and termination behavior. Specifically, a sliding-window semantic similarity metric penalizes low-progress reasoning segments, while a normalized external-redundancy metric discourages post-answer continuation. Experiments on GSM8K, MATH500, and AIME24 across different model scales show that our method reduces average reasoning length by 41.3% on the 1.5B model and 40.1% on the 7B model, while preserving competitive accuracy and achieving the best overall accuracy-efficiency score among evaluated baselines. The learned compression behavior further transfers to out-of-domain reasoning tasks, including GPQA and LiveCodeBench. More importantly, our analysis reveals a clear asymmetry between the two redundancy types: external redundancy can be largely removed with little performance loss, whereas internal redundancy compression follows a sensitive accuracy-efficiency trade-off. These results suggest that effective CoT compression should optimize semantic efficiency rather than sequence length alone, offering a principled route toward more concise, efficient, and interpretable LRMs.

cs.AI