SearcharxivSearch

arXiv subjects

Zipei Zhang

Publications and source records attributed to Zipei Zhang.

8 recordsLinked to original sources

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding. This capability is essential for real-world applications such as intelligent video editing, surveillance, and human-computer interaction. However, existing datasets are largely limited to simple or homogeneous audio-visual scenes with coarse annotations. Such oversimplified settings bias evaluation toward static audio-visual co-occurrence, rather than rigorously assessing robust spatiotemporal modeling and cross-modal reasoning in complex, dynamic scenes. To address these limitations, we introduce AVTrack, a human-centric audio-visual instance segmentation (AVIS) dataset designed for dynamic real-world scenarios. AVTrack features diverse and challenging conditions, including camera motion, visual occlusions, and position changes. Evaluations of representative AVIS methods on AVTrack reveal substantial performance degradation, establishing AVTrack as a challenging benchmark for robust human-centric audio-visual scene understanding in complex environments. We further provide a simple yet effective baseline to facilitate future research. Project website: https://FudanCVL.github.io/AVTrack/

cs.CV

AgentMark: Utility-Preserving Behavioral Watermarking for Agents

LLM-based agents are increasingly deployed to autonomously solve complex tasks, raising urgent needs for IP protection and regulatory provenance. While content watermarking effectively attributes LLM-generated outputs, it fails to directly identify the high-level planning behaviors (e.g., tool and subgoal choices) that govern multi-step execution. Critically, watermarking at the planning-behavior layer faces unique challenges: minor distributional deviations in decision-making can compound during long-term agent operation, degrading utility, and many agents operate as black boxes that are difficult to intervene in directly. To bridge this gap, we propose AgentMark, a behavioral watermarking framework that embeds multi-bit identifiers into planning decisions while preserving utility. It operates by eliciting an explicit behavior distribution from the agent and applying distribution-preserving conditional sampling, enabling deployment under black-box APIs while remaining compatible with action-layer content watermarking. Experiments across embodied, tool-use, and social environments demonstrate practical multi-bit capacity, robust recovery from partial logs, and utility preservation. The code is available at https://github.com/Tooooa/AgentMark.

cs.CR

On Gauging Finite Symmetries by Higher Gauging Condensation Defects

Based on the work by C{\'o}rdova-Costa-Hsin (arXiv:2412.16681), we propose an EFT-style, Lagrangian procedure to gauge finite 0-form symmetries in untwisted Dijkgraaf-Witten gauge theories on closed oriented manifolds using higher gauging condensation defects and point out its limitations. Using this proposal, we construct effective actions of untwisted Dijkgraaf-Witten theories with Heisenberg gauge group over $\mathbb{Z}_p$ and show that the braiding data from Hopf link and the fusion rules match with the expected discrete gauge theories. We also study the symTFT implications of these effective Lagrangians and clarify their relations with higher group global symmetries.

hep-th

GSDFuse: Capturing Cognitive Inconsistencies from Multi-Dimensional Weak Signals in Social Media Steganalysis

The ubiquity of social media platforms facilitates malicious linguistic steganography, posing significant security risks. Steganalysis is profoundly hindered by the challenge of identifying subtle cognitive inconsistencies arising from textual fragmentation and complex dialogue structures, and the difficulty in achieving robust aggregation of multi-dimensional weak signals, especially given extreme steganographic sparsity and sophisticated steganography. These core detection difficulties are compounded by significant data imbalance. This paper introduces GSDFuse, a novel method designed to systematically overcome these obstacles. GSDFuse employs a holistic approach, synergistically integrating hierarchical multi-modal feature engineering to capture diverse signals, strategic data augmentation to address sparsity, adaptive evidence fusion to intelligently aggregate weak signals, and discriminative embedding learning to enhance sensitivity to subtle inconsistencies. Experiments on social media datasets demonstrate GSDFuse's state-of-the-art (SOTA) performance in identifying sophisticated steganography within complex dialogue environments. The source code for GSDFuse is available at https://github.com/NebulaEmmaZh/GSDFuse.

cs.CR

Agent Guide: A Simple Agent Behavioral Watermarking Framework

The increasing deployment of intelligent agents in digital ecosystems, such as social media platforms, has raised significant concerns about traceability and accountability, particularly in cybersecurity and digital content protection. Traditional large language model (LLM) watermarking techniques, which rely on token-level manipulations, are ill-suited for agents due to the challenges of behavior tokenization and information loss during behavior-to-action translation. To address these issues, we propose Agent Guide, a novel behavioral watermarking framework that embeds watermarks by guiding the agent's high-level decisions (behavior) through probability biases, while preserving the naturalness of specific executions (action). Our approach decouples agent behavior into two levels, behavior (e.g., choosing to bookmark) and action (e.g., bookmarking with specific tags), and applies watermark-guided biases to the behavior probability distribution. We employ a z-statistic-based statistical analysis to detect the watermark, ensuring reliable extraction over multiple rounds. Experiments in a social media scenario with diverse agent profiles demonstrate that Agent Guide achieves effective watermark detection with a low false positive rate. Our framework provides a practical and robust solution for agent watermarking, with applications in identifying malicious agents and protecting proprietary agent systems.

cs.AI

SymSETs and self-dualities under gauging non-invertible symmetries

The self-duality defects under discrete gauging in a categorical symmetry $\mathcal{C}$ can be classified by inequivalent ways of enriching the bulk SymTFT of $\mathcal{C}$ with $\mathbb{Z}_2$ 0-form symmetry. The resulting Symmetry Enriched Topological (SET) orders will be referred to as $\textit{SymSETs}$ and are parameterized by choices of $\mathbb{Z}_2$ symmetries, as well as symmetry fractionalization classes and discrete torsions. In this work, we consider self-dualities under gauging $\textit{non-invertible}$ $0$-form symmetries in $2$-dim QFTs and explore their SymSETs. Unlike the simpler case of self-dualities under gauging finite Abelian groups, the SymSETs here generally admit multiple choices of fractionalization classes. We provide a direct construction of the SymSET from a given duality defect using its $\textit{relative center}$. Using the SymSET, we show explicitly that changing fractionalization classes can change fusion rules of the duality defect besides its $F$-symbols. We consider three concrete examples: the maximal gauging of $\operatorname{Rep} H_8$, the non-maximal gauging of the duality defect $\mathcal{N}$ in $\operatorname{Rep} H_8$ and $\operatorname{Rep} D_8$ respectively. The latter two cases each result in 6 fusion categories with two types of fusion rules related by changing fractionalization class. In particular, two self-dualities of $\operatorname{Rep} D_8$ related by changing the fractionalization class lead to $\operatorname{Rep} D_{16}$ and $\operatorname{Rep} SD_{16}$ respectively. Finally, we study the physical implications such as the spin selection rules and the SPT phases for the aforementioned categories.

hep-th

Exploring $G$-ality defects in 2-dim QFTs

The Tambara-Yamagami (TY) fusion category symmetry $\text{TY}(\mathbb{A},\chi,\epsilon)$ describes the enhanced non-invertible self-duality symmetry of a $2$-dim QFT under gauging a finite Abelian group $\mathbb{A}$. We generalize the enhanced non-invertible symmetries by considering twisted gauging which allows stacking $\mathbb{A}$-SPTs before and after the gauging. Such non-invertible symmetries can be obtained from invertible anyon permutation symmetries of the $3$-dim SymTFT. Consider a finite group $G$ formed by (un)twisted gaugings of $\mathbb{A}$, a $2$-dim QFT invariant under topological manipulations in $G$ admits non-invertible \textit{$G$-ality defects}. We study the classification and the physical implication of the $G$-ality defects using the SymTFT and the group-theoretical fusion categories, with three concrete examples. 1) Triality with $\mathbb{A} = \mathbb{Z}_N \times \mathbb{Z}_N$ where $N$ is coprime with $3$. The classification was previously determined by Jordan and Larson where the data is similar to the $\text{TY}$ fusion categories, and we determine the anomaly of these fusion categories. 2) $p$-ality with $\mathbb{A} = \mathbb{Z}_p \times \mathbb{Z}_p$ where $p$ is an odd prime. We consider two such categories $\mathcal{P}_{\pm,m}$ which are distinguished by different choices of the symmetry fractionalization, a new data that does not appear in the TY classification, and show that they have distinct anomaly structures and spin selection rules. 3) $S_3$-ality with $\mathbb{A} = \mathbb{Z}_N \times \mathbb{Z}_N$. We study their classification explicitly for $N < 20$ via SymTFT, and provide a group-theoretical construction for certain $N$. We find $N=5$ is the minimal $N$ to admit an $S_3$-ality and $N=11$ is the minimal $N$ to admit a group-theoretical $S_3$-ality.

hep-th

De Sitter Quantum Loops as the origin of Primordial Non-Gaussianities

It was pointed out recently that in some inflationary models quantum loops containing a scalar of mass $m$ that couples to the inflaton can be the dominant source of primordial non-Gaussianities. We explore this phenomenon in the simplest such model focusing on the behavior of the primordial curvature fluctuations for small $m/H$. Explicit calculations are done for the three and four point curvature fluctuation correlations. Constraints on the parameters of the model from the CMB limits on primordial non-Gaussianity are discussed. The bi-spectrum in the squeezed limit and the tri-spectrum in the compressed limit are examined. The form of the $n$-point correlations as any partial sum of wave vectors get small is determined.

hep-ph