SearcharxivSearch

arXiv subjects

Wenyuan Yang

Publications and source records attributed to Wenyuan Yang.

At least 19 recordsLinked to original sources

Growth Gaps for Quotients by Confined Subgroups

In this paper, we establish growth gaps for quotients by confined subgroups in groups admitting a statistically convex-cocompact action with contracting elements. We also discuss applications to uniformly recurrent subgroups.

math.GR

Sublinear projection tracking in acylindrically hyperbolic groups

We study projection phenomena in word metrics of finitely generated acylindrically hyperbolic groups. For a loxodromic WPD element acting on a hyperbolic space, we prove that shortest projection in the word metric to the corresponding cyclic subgroup sublinearly tracks the pullback of shortest projection to its axis in the hyperbolic space. As applications, we obtain effective upper bounds for growth functions and construct proper quotients whose growth rates converge to that of the original group. We further prove a growth--cogrowth inequality for confined subgroups in both acylindrically hyperbolic groups and Morse local-to-global groups with Morse elements.

math.GR

Growth gaps and exponential genericity in acylindrically hyperbolic groups

We prove that, for every finite generating set of an acylindrically hyperbolic group, the set of non-WPD elements has strictly smaller exponential growth rate. Equivalently, WPD elements are exponentially generic. As applications, we prove growth tightness and cogrowth tightness for acylindrically hyperbolic groups.

math.GR

Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation

As LLMs become increasingly woven into everyday workflows, user queries sent to cloud hosted LLMs routinely mix task-essential content with task non-essential sensitive disclosures, yet type based PII redaction is context agnostic and may raise two issues: over disclosing untyped sensitive context and over removing answer bearing spans. We recast privacy preserving query rewriting under Contextual Integrity: a span should be forwarded only if it is necessary for the task. We introduce DelegateCI-Bench, the first task based Contextual Integrity benchmark for privacy-conscious delegation, comprising 3,167 samples that combine high quality synthetic data spanning 11 tasks and 20 task types, WildChat based real user queries, and a medical challenge set with dense sensitive information. Building on this benchmark, we propose a CI-guided reinforcement learning framework that converts essential and non-essential sensitive spans into verifiable optimization signals, and train a query rewriter to preserve task critical information while suppressing unnecessary sensitive disclosure. Experiments show that our learned rewriter achieves the best privacy-utility tradeoff, achieving up to +10.1 average utility over on-device baselines.

cs.CR

SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking

Large-scale vision-language models, especially CLIP, have demonstrated remarkable performance across diverse downstream tasks. Soft prompts, as carefully crafted modules that efficiently adapt vision-language models to specific tasks, necessitate effective copyright protection. In this paper, we investigate model copyright protection by auditing whether suspicious third-party models incorporate protected soft prompts. While this can be viewed as a special case of model ownership auditing, our analysis shows that existing techniques are ineffective due to prompt learning's unique characteristics. Non-intrusive auditing is inherently prone to false positives when independent models share similar data distributions with victim models. Intrusive approaches also fail: backdoor methods designed for CLIP cannot embed functional triggers, while extending traditional DNN backdoor techniques to prompt learning suffers from harmfulness and ambiguity challenges. We find that these failures in intrusive auditing stem from the same fundamental reason: watermarking operates within the same decision space as the primary task yet pursues opposing objectives. Motivated by these findings, we propose sequential watermarking for soft prompts (SWAP), which implants watermarks into a different and more complex space. SWAP encodes watermarks through a specific order of defender-specified out-of-distribution classes, inspired by the zero-shot prediction capability of CLIP. This watermark, which is embedded in a more complex space, keeps the original prediction label unchanged, making it less opposed to the primary task. We further design a hypothesis-test-guided verification protocol for SWAP and provide a theoretical analysis of when verification works. Extensive experiments on 11 datasets demonstrate SWAP's effectiveness, harmlessness, and robustness against potential attacks.

cs.CR

Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses

Image steganography is widely used to protect user privacy and enable covert communication. However, it can also be abused by the adversary as a covert channel to bypass content moderation, disseminate harmful semantics, and even hide malicious instructions in images to elicit dangerous outputs from large models, posing a practical security risk that continues to evolve. To address the lack of a unified and systematic evaluation framework, we propose SADBench, a systematic benchmark that assesses the adversary's ability to inject harmful secrets via steganography and the defender's ability to detect such threats through steganalysis. Crucially, SADBench comprises $4$ core tasks, namely steganography attack capability evaluation, steganalysis defense capability evaluation, efficiency evaluation, and transferability evaluation. It evaluates both image-payload and text-payload steganography across diverse cover distributions, utilizing harmful visual semantics and toxic instructions to simulate malicious attacks. Across a broad set of attacks and detectors, SADBench reveals that (i) INN and autoencoder-based methods demonstrate superior stability compared to other architectures, (ii) in-domain detection is near-perfect and cheaper than generation, (iii) a critical asymmetry exists in transferability where attacks robustly generalize to new distributions while detectors fail to adapt, and (iv) real-world threats persist on social media, where payloads either survive minimal compression or effectively adapt to aggressive compression via simulated training. Overall, SADBench establishes a systematic, reproducible, and extensible framework to quantify risks, paving the way for measurable and security-driven advancements in steganography defense.

cs.CR

Missing Pattern Tree based Decision Grouping and Ensemble for Enhancing Pair Utilization in Deep Incomplete Multi-View Clustering

Real-world multi-view data often exhibit highly inconsistent missing patterns, posing significant challenges for incomplete multi-view clustering (IMVC). Although existing IMVC methods have made progress from both imputation-based and imputation-free routes, they largely overlook the issue of pair underutilization. Specifically, inconsistent missing patterns prevent incomplete but available multi-view pairs from being fully exploited, thereby limiting the model performance. To address this limitation, we propose a novel missing-pattern tree based IMVC framework. Specifically, to fully leverage available multi-view pairs, we first introduce a missing-pattern tree model to group data into multiple decision sets according to their missing patterns, and then perform multi-view clustering within each set. Furthermore, a multi-view decision ensemble module is proposed to aggregate clustering results across all decision sets. This module infers uncertainty-based weights to suppress unreliable clustering decisions and produce robust outputs. Finally, we develop an ensemble-to-individual knowledge distillation module module, which transfers ensemble knowledge to view-specific clustering models. This design enables mutual enhancement between ensemble and individual modules by optimizing cross-view consistency and inter-cluster discrimination losses. Extensive theoretical analysis supports our key designs, and empirical experiments on multiple benchmark datasets demonstrate that our method effectively mitigates the pair underutilization issue and achieve superior IMVC performance.

cs.LG

Two Characterizations of Geometrically Infinite Actions on Gromov Hyperbolic Spaces

We provide two new characterizations of geometrically infinite actions on Gromov hyperbolic spaces: one in terms of the existence of escaping geodesics, and the other via the presence of uncountably many non-conical limit points. These results extend corresponding theorems of Bonahon, Bishop, and Kapovich--Liu from the settings of Kleinian groups and pinched negatively curved manifolds to discrete groups acting properly on proper Gromov hyperbolic spaces.

math.GR

Boundary actions of Bass-Serre Trees and the applications to $C^*$-algebras

In this paper, we study Bass-Serre theory from the perspectives of $C^*$-algebras and topological dynamics. In particular, we investigate the actions of fundamental groups of graphs of groups on their Bass-Serre trees and the associated boundaries, through which we identify new families of $C^*$-simple groups including certain tubular groups, fundamental groups of certain graphs of groups with one vertex group acylindrically hyperbolic and outer automorphism groups $\operatorname{Out}(BS(p, q))$ of Baumslag-Solitar groups. In addition, we study $n$-dimensional Generalized Baumslag-Solitar ($\text{GBS}_n$) groups. We first recover a result by Minasyan and Valiunas on the characterization of $C^*$-simplicity for $\text{GBS}_1$ groups and identify new $C^*$-simple $\text{GBS}_n$ groups including the Leary-Minasyan group. These $C^*$-simple groups also provide new examples of $C^*$-selfless groups and highly transitive groups. Moreover, we demonstrate that natural boundary actions of these $C^*$-simple fundamental groups of graphs of groups give rise to the new purely infinite crossed product $C^*$-algebras.

math.OA

An extreme boundary of acylindrically hyperbolic groups

We prove that every acylindrically hyperbolic group admits a minimal and extremely proximal action on a compact metrizable space. If there are no nontrivial finite normal subgroups, then the action is topologically free. This answers positively a question of Ozawa and the applications to $C^\ast$-algebras are discussed.

math.GR

Global-Graph Guided and Local-Graph Weighted Contrastive Learning for Unified Clustering on Incomplete and Noise Multi-View Data

Recently, contrastive learning (CL) plays an important role in exploring complementary information for multi-view clustering (MVC) and has attracted increasing attention. Nevertheless, real-world multi-view data suffer from data incompleteness or noise, resulting in rare-paired samples or mis-paired samples which significantly challenges the effectiveness of CL-based MVC. That is, rare-paired issue prevents MVC from extracting sufficient multi-view complementary information, and mis-paired issue causes contrastive learning to optimize the model in the wrong direction. To address these issues, we propose a unified CL-based MVC framework for enhancing clustering effectiveness on incomplete and noise multi-view data. First, to overcome the rare-paired issue, we design a global-graph guided contrastive learning, where all view samples construct a global-view affinity graph to form new sample pairs for fully exploring complementary information. Second, to mitigate the mis-paired issue, we propose a local-graph weighted contrastive learning, which leverages local neighbors to generate pair-wise weights to adaptively strength or weaken the pair-wise contrastive learning. Our method is imputation-free and can be integrated into a unified global-local graph-guided contrastive learning framework. Extensive experiments on both incomplete and noise settings of multi-view data demonstrate that our method achieves superior performance compared with state-of-the-art approaches.

cs.LG

SoK: Large Language Model Copyright Auditing via Fingerprinting

The broad capabilities and substantial resources required to train Large Language Models (LLMs) make them valuable intellectual property, yet they remain vulnerable to copyright infringement, such as unauthorized use and model theft. LLM fingerprinting, a non-intrusive technique that compares the distinctive features (i.e., fingerprint) of LLMs to identify whether an LLM is derived from another, offers a promising solution to copyright auditing. However, its reliability remains uncertain due to the prevalence of diverse model modifications and the lack of standardized evaluation. In this SoK, we present the first comprehensive study of the emerging LLM fingerprinting. We introduce a unified framework and taxonomy that structures the field: white-box methods are classified based on their feature source as static, forward-pass, or backward-pass fingerprinting, while black-box methods are distinguished by their query strategy as either untargeted or targeted. Furthermore, we propose LeaFBench, the first systematic benchmark for evaluating LLM fingerprinting under realistic deployment scenarios. Built upon 7 mainstream foundation models and comprising 149 distinct model instances, LeaFBench integrates 13 representative post-development techniques, spanning both parameter-altering methods (e.g., fine-tuning, quantization) and parameter-independent techniques (e.g., system prompts, RAG). Extensive experiments on LeaFBench reveal the strengths and weaknesses of existing methods, thereby outlining future research directions and critical open problems in this emerging field. The code is available at https://github.com/shaoshuo-ss/LeaFBench.

cs.CR

Boundary actions of CAT(0) spaces: topological freeness and applications to $C^*$-algebras

In this paper, we study topological dynamics on the visual boundary and several combinatorial boundaries associated to $\operatorname{CAT}(0)$ spaces. Through verifying the freeness of Myrberg points on the boundaries, we prove that a large class of these boundary actions are topologically free strong boundary actions. These include certain visual boundary actions obtained from proper isometric actions of groups on proper $\operatorname{CAT}(0)$ spaces with rank-one elements, horofunction boundary actions from actions of irreducible finitely generated infinite non-affine Coxeter groups on the Caylay graphs, and Roller-type boundary actions from certain group actions on irreducible $\operatorname{CAT}(0)$ cube complexes. This in particular leads to a new proof of Kar-Sageev's topological freeness result for Roller boundary actions of $\operatorname{CAT}(0)$ cube complexes and generalizes Klisee's topological freeness result on horofunction boundaries from hyperbolic and right angled Coxeter groups to the general case. As applications to $C^*$-algebras, our work yields new examples of $C^*$-selfless groups and of exact, purely infinite, simple reduced crossed product $C^\ast$-algebras.

math.GR

Ancona inequalities along generic geodesic rays

This paper presents several versions of the Ancona inequality for finitely supported, irreducible random walks on non-amenable groups. We first study a class of Morse subsets with narrow points and prove the Ancona inequality around these points in any finitely generated non-amenable group. This result implies the inequality along all Morse geodesics and recovers the known case for relatively hyperbolic groups. We then consider any geometric action of a non-amenable group with contracting elements. For such groups, we construct a class of generic geodesic rays, termed proportionally contracting rays, and establish the Ancona inequality along a sequence of good points. This leads to an embedding of a full-measure subset of the horofunction boundary into the minimal Martin boundary. A stronger Ancona inequality is established for groups acting geometrically on an irreducible CAT(0) cube complex with a Morse hyperplane. In this setting, we show that the orbital maps extend continuously to a partial boundary map from a full-measure subset of the Roller boundary into the minimal Martin boundary. Finally, we provide explicit examples, including right-angled Coxeter groups (RACGs) defined by an irreducible graph with at least one vertex not belonging to any induced 4-cycle.

math.GR

Hausdorff Dimension of non-conical and Myrberg limit sets

In this paper, we develop techniques to study the Hausdorff dimensions of non-conical and Myrberg limit sets for groups acting on negatively curved spaces. We establish maximality of the Hausdorff dimension of the non-conical limit set of $G$ in the following cases. 1. $M$ is a finite volume complete Riemannian manifold of pinched negative curvature and $G$ is an infinite normal subgroups of infinite index in $π_1(M)$. 2. $G$ acts on a regular tree $X$ with $X/G$ infinite and amenable (dimension 1). 3. $G$ acts on the hyperbolic plane $\mathbb H^2$ such that $\mathbb H^2/G$ has Cheeger constant zero (dimension 2). 4. $G$ is a finitely generated geometrically infinite Kleinian group (dimension 3). We also show that the Hausdorff dimension of the Myrberg limit set is the same as the critical exponent, confirming a conjecture of Falk-Matsuzaki.

math.GR

ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors

Multilingual audio-text retrieval (ML-ATR) is a challenging task that aims to retrieve audio clips or multilingual texts from databases. However, existing ML-ATR schemes suffer from inconsistencies for instance similarity matching across languages. We theoretically analyze the inconsistency in terms of both multilingual modal alignment direction error and weight error, and propose the theoretical weight error upper bound for quantifying the inconsistency. Based on the analysis of the weight error upper bound, we find that the inconsistency problem stems from the data distribution error caused by random sampling of languages. We propose a consistent ML-ATR scheme using 1-to-k contrastive learning and audio-English co-anchor contrastive learning, aiming to mitigate the negative impact of data distribution error on recall and consistency in ML-ATR. Experimental results on the translated AudioCaps and Clotho datasets show that our scheme achieves state-of-the-art performance on recall and consistency metrics for eight mainstream languages, including English. Our code will be available at https://github.com/ATRI-ACL/ATRI-ACL.

cs.SD

Marked length spectrum rigidity in groups with contracting elements

This paper presents a study of the well-known marked length spectrum rigidity problem in the coarse-geometric setting. For any two (possibly non-proper) group actions $G\curvearrowright X_1$ and $G\curvearrowright X_2$ with contracting property, we prove that if the two actions have the same marked length spectrum, then the orbit map $Go_1\to Go_2$ must be a rough isometry. In the special case of cusp-uniform actions, the rough isometry can be extended to the entire space. This generalizes the existing results in hyperbolic groups and relatively hyperbolic groups. In addition, we prove a finer marked length spectrum rigidity from confined subgroups and further, geometrically dense subgroups. Our proof is based on the Extension Lemma and uses purely elementary metric geometry. This study produces new results and recovers existing ones for many more interesting groups through a unified and elementary approach.

math.GR