SearcharxivSearch

arXiv subjects

Liang Yang

Publications and source records attributed to Liang Yang.

At least 19 recordsLinked to original sources

Resource Allocation for Secure Dual-UAV-Assisted ISAC System

Integrated sensing and communication (ISAC) is a rising technology in the next wireless communication networks, enabling the simultaneous execution of communication and sensing tasks by fully utilizing limited spectrum resources. In this work, we investigate the secrecy performance of a dual-uncrewed aerial vehicle (UAV)-assisted secure ISAC system. Specifically, a base station UAV communicates with users and transmits radar signals to locate potential eavesdroppers, while simultaneously providing information to a jammer UAV to perform jamming tasks. Considering constraints such as maximum UAV velocity, transmit power, propulsion energy, and sensing thresholds, we maximize the average secrecy rate by optimizing user scheduling strategies, time allocation, transmit power, and UAV trajectories. The presence of a non-convex problem, originating from tightly coupled variables, is tackled by an efficient iterative algorithm. In particular, the original optimization problem is decomposed into six subproblems, and non-convex subproblems are transformed into approximately convex forms via successive convex approximation. Then, block coordinate descent techniques are employed to solve all subproblems sequentially. Numerical results demonstrate the convergence and effectiveness of the proposed algorithm.

cs.IT

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.

cs.CL

DRL-Based Secure Transmission for Rotatable Antenna-Enabled Low-Altitude ISAC Systems

The development of the low-altitude economy has driven innovation in intelligent antenna systems within ISAC systems. In this paper, we investigate a Rotatable Antenna (RA)-enabled low-altitude integrated sensing and communication (ISAC) system. In practical terms, the RA array can flexibly adjust the three-dimensional (3D) beam direction of each antenna to enhance array directional gain, thereby improving the communication security of legitimate mobile users against potential eavesdropping risks from the unmanned aerial vehicle (UAV). Our objective is to maximize the minimum secrecy rate (SR) by jointly optimizing transmit beamforming matrix, transmit and receive RAs' pointing matrices. To this end, an multi-agent proximal policy optimization with three improvement mechanisms (MAPPO-T) algorithm is proposed to cope with the issue of complex multi-agent collaborative decision-making problem. Simulation results show that the introduction of RAs can effectively improve SR performance compared to the traditional fixed orientation antenna (FOA)-based system. In addition, the proposed MAPPO-T algorithm validate the superiority compared to the standard MAPPO algorithm.

eess.SP

Secure Relay Low-Altitude Networks via Hybrid Fixed-Position and Rotatable Antenna Arrays

In this paper, a relay network with hybrid fixed-position and rotatable antenna arrays is proposed. The deployment of rotatable arrays in conventional relay networks is considered to provide more secure communications for low-altitude economy applications. Specifically, both the base station and the relay station are equipped with fixed-position antenna arrays and rotatable arrays to serve ground users and aerial users, respectively. To address the challenge of multi-user interference, a low-cost reconfigurable intelligent surface is exploited as a candidate path. Accordingly, under constraints on transmit power, user quality of service, rotatable range, and path selection, the objective is to maximize the worst-case secrecy rate (SR) through joint beamforming, power allocation, and rotatable antenna orientation design. First, the SR performance in the single-user scenario is investigated, and a step-by-step leakage-based scheme is proposed. Then, the general multi-user scenario is studied, and a Distributional Soft Actor-Critic with Three refinements (DSAC-T)-based learning scheme, which supports hybrid discrete and continuous actions, is proposed to maximize the worst-case SR. Simulation results validate the effectiveness of the proposed schemes. The proposed schemes achieve approximately a twofold improvement in SR performance compared to isotropic antennas. The proposed system achieves approximately 71.4\% power saving, 55\% antenna saving, and can serve more users.

eess.SP

Trade-off for Secure UAV-ISCC Systems

The integrated sensing, communication, and computing (ISCC) system overcomes the limitations of conventional standalone architectures. Through resource sharing and collaborative design, it dynamically optimizes and jointly enhances communication, sensing, and computing performance, thereby significantly improving overall system efficiency. This work investigates the performance trade-off among secure communication rate, radar estimation rate, and computational energy efficiency in an uncrewed aerial vehicle (UAV)-assisted ISCC system. By jointly optimizing the UAV's three-dimensional (3D) trajectory, beamforming, user scheduling, and computational frequency, three optimization problems are formulated to maximize the average secrecy rate, sensing rate, and computational energy efficiency, respectively, thus establishing the system's performance boundaries under diverse scenarios. On this basis, the trade-off among security, sensing, and computation is further explored with the goal of maximizing the normalized weighted sum of the three performance metrics, which provides a theoretical basis for the performance-coordinated design of aerial ISCC systems.

cs.IT

A five-variable counterexample to the Hessian conjecture, and the low-dimensional status of the Jacobian and Hessian conjectures

We exhibit an explicit integer polynomial in five variables, of total degree $14$ and with constant Hessian determinant $128$, whose gradient is not injective. Consequently its formal Legendre transform is not a polynomial, and the Hessian conjecture $\HC_5$ is false. The counterexample is obtained from the six-variable doubling of Alp\"oge's 2026 Jacobian counterexample by a one-variable \emph{Schur descent}---a partial Legendre transform in a single variable. Combined with de~Bondt's theorem that $\HC_n$ holds for $n\le3$, with the elementary doubling and stabilization bridges relating the Jacobian conjectures $\JC_n$ to the Hessian conjectures $\HC_n$, and with Alp\"oge's refutation of $\JC_3$, this decides the Hessian conjecture in every dimension except $n=4$: $\HC_n$ is true for $n\le3$, false for $n\ge5$, and open only at $n=4$. Exactly two statements of the two families remain unsettled, $\JC_2$ and $\HC_4$, linked by $\HC_4 \Rightarrow \JC_2$. Along the way we record, as a warm-up, an explicit six-variable counterexample to $\HC_6$ with constant Hessian determinant $-4$ and non-injective gradient. This note adds the five-variable counterexample to, and updates the status recorded in, the first author's earlier educational preprint \cite{MengRG2026}.

math.AC

Homological rigidity of quiver representations over $\mathbb{F}_1$

We establish a homological rigidity phenomenon for the category of representations of quivers over the virtual field $\mathbb{F}_1$, which is inherently non-additive and does not admit classical homological algebra tools. We prove that all higher Yoneda extension groups vanish beyond degree two for arbitrary quivers, including infinite ones. Consequently, the global dimension of the category is universally bounded by 2. Moreover, we obtain a complete classification of quivers according to their homological dimension, which is determined solely by the underlying orientation structure.

math.RT

Shift-MoE-Based DJSCC for CSI Feedback in Multi-User Pinching-Antenna Systems

In frequency-division duplexing systems, the performance gains of pinching-antenna systems (PASS) critically depend on accurate channel state information (CSI) at the base station. However, PASS CSI exhibits structured correlations over the waveguide-antenna grid and pronounced heterogeneity across users, making conventional fixed feedback mappings difficult to generalize. To address this challenge, this letter proposes an end-to-end CSI feedback scheme over a noisy uplink feedback link based on deep joint source-channel coding, termed Shift-based Mixture-of-Experts (Shift-MoE). Specifically, Shift-MoE leverages channel-grouped one-step shift operations to capture grid dependencies without global attention, and employs a gated multilayer perceptron mixture-of-experts module to adapt to heterogeneous CSI statistics across users. Numerical results demonstrate that the proposed Shift-MoE consistently outperforms representative learning-based CSI feedback baselines in normalized mean squared error and remains effective under different system parameter settings.

eess.SP

Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos

Hateful videos have become prevalent on online platforms, highlighting an urgent need for effective detection. However, existing studies primarily focus on binary classification and fail to provide contextual rationales that reveal the implicit meanings behind these judgments, significantly undermining model explainability. To fill this gap, we aim to achieve explainable hateful video detection, enabling models to provide contextual rationales that integrate relevant evidence and logical reasoning alongside decisions. This approach can comprehensively enhance the understanding of video content and the explainability of the decision-making process. We first introduce two datasets, Ex-HateMM and Ex-ImpliHateVid, for explainable hateful video detection. Each dataset provides fine-grained annotations of multimodal harmful elements, along with contextual rationales. We then propose an Information Augmentation and Reasoning Enhancement (IARE) framework designed for explainable detection. The framework employs an information augmentation phase that leverages the multimodal chain-of-thought to integrate harmful elements, thereby enriching rationale evidence. Additionally, IARE incorporates a reasoning enhancement phase, in which Direct Preference Optimization guides the model toward correct reasoning paths and away from incorrect ones, thereby improving the logical coherence of its justifications. We conduct extensive experiments on the two datasets, comparing multiple baselines with our proposed IARE framework. The results demonstrate that IARE achieves state-of-the-art performance while also generating accurate rationales.

cs.CL

A Graph Foundation Model with Spectral Parsing and Prototype-Guided Spatial Propagation

Graph foundation models aim to learn transferable knowledge from diverse graphs for generalization to unseen graphs and tasks. Unlike text and images, graphs lack a shared vocabulary or regular spatial grid, making cross-graph transfer challenging. This challenge comes from both feature discrepancies and, more critically, diverse graph structures. Existing GFMs mainly improve transferability by unifying feature spaces or incorporating structural tokens and vocabularies. However, existing topology-aware designs still have limitations. Structural tokens are usually discrete, while structural vocabularies often rely on predefined substructures such as trees and cycles, whose limited coverage may miss richer relational patterns across graphs. Moreover, graph signals contain both high-frequency local patterns and smoother low-frequency patterns, which require different propagation behaviors. These components are often entangled in raw graph signals, while this spectral perspective is rarely explored in existing GFMs. To address these challenges, we propose SPG, a graph foundation model with spectral parsing and prototype-guided spatial propagation. SPG applies learnable Chebyshev filters to decompose node features into multiple spectral responses, reducing the mismatch between frequency-specific graph signals and propagation behaviors. It then constructs a Gromov-Wasserstein prototype geometry to distill transferable pairwise relations beyond predefined substructures into a shared structural space. The learned prototype geometry is further projected back as a prototype-guided propagation operator. Experiments demonstrate consistent improvements in cross-domain generalization.

cs.LG

On Secure EKF-enhanced UAV-ISAC Systems

Integrated sensing and communication (ISAC) has emerged as a promising key technology for future wireless networks, enabling the efficient coordination of sensing and communication functions within limited resources. This work investigates a secure ISAC system assisted by an uncrewed aerial vehicle (UAV). By incorporating the extended Kalman filter (EKF), the proposed system is capable of delivering communication services to legitimate users while simultaneously jamming eavesdroppers and performing joint prediction and tracking of the trajectories of both legitimate and illegitimate users. Considering practical constraints such as {sensing beamwidth}, transmit power, and UAV's propulsion energy consumption, the secrecy rate is maximized through the joint design of transmit beamforming and UAV trajectory. To tackle the resulting highly non-convex optimization problem, an efficient iterative algorithm is developed by integrating block coordinate descent, successive convex approximation, and EKF, thereby yielding a high-quality suboptimal solution. Extensive simulation results validate the superior performance of the proposed scheme compared to benchmarks.

cs.IT

On higher extensions of quiver representations over $\mathbb{F}_1$

We show that higher extension spaces between finite-dimensional nilpotent $\mathbb{F}_1$-representations maybe infinite-dimensional, thereby clarifying a misconception in the literature. Our examples arise from cyclic quivers. In particular, for a cyclic quiver $\Delta_n$, we show that $\operatorname{Ext}^3(-,-)$ vanishes for any pair of finite-dimensional nilpotent $\mathbb{F}_1$-representations of $\Delta_n$, while $\operatorname{Ext}^2(-,-)$ is infinite-dimensional for any pair of simple representations.

math.RT

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imply faithful use of visual evidence. Prior studies have identified three shortcuts that inflate benchmark performance. First, linguistic priors and lexical cues in questions often enable models to infer plausible answers without seeing the image. Second, coarse global semantics from the visual encoder can bypass fine-grained local details. Third, in some ``think-with-images'' benchmarks, corrupting the intermediate images returned by visual tools barely affects the final answer. These findings suggest that higher input resolution or larger question pools alone do not elicit genuine active visual search. To address this, we introduce VisualNeedle, a challenging, information-dense, and fine-grained benchmark for scenes where critical evidence is spatially constrained to minute regions and not discernible at a glance. We further propose a counterfactual crop-black setting, which replaces crops returned by tools with black images of the same size, to test whether tool-enabled performance truly relies on intermediate visual evidence. We evaluate 9 promninent MLLMs across three settings: no-tool, standard tool-enabled, and crop-black. No-tool accuracy stays below 20\%, and the best tool-enabled model reaches only 56.01\%, still trailing the 63.00% human majority-vote accuracy. These results reveal persistent limitations in fine-grained visual search, while the crop-black ablation confirms that success on VisualNeedle hinges on genuine intermediate visual evidence.

cs.CV

Distinguishing Right from Wrong in Debates: Attribution Analysis of Chinese Harmful Memes

Research on harmful meme detection has garnered significant attention, resulting in the development of numerous datasets and methods. However, progress in detecting Chinese harmful memes lags considerably, primarily due to two challenges: first, accurately assessing a meme's harmfulness depends heavily on understanding deep cultural context; second, many memes are semantically ambiguous, making harmfulness highly subjective. To address these issues, we focus on the interpretable detection of Chinese harmful memes by constructing the first Chinese harmful meme explanation dataset, Ex-ToxiCN-MM. This dataset offers opposing interpretations, categorized as "harmful" and "non-harmful", for each meme, aiming to rigorously evaluate a model's ability to discern and comprehend ambiguous, culturally grounded content. We built a specialized knowledge base of Chinese cultural concepts and offensive vocabulary to supply models with essential prior knowledge (C-HarmKB). To address the ambiguity and lack of background knowledge in meme attribution, we have developed a comprehensive attribution analysis framework, RIKE, which includes an Attribution Knowledge Enhancement module (AKE) and a Relative Intent Reasoning module (RIR). Extensive quantitative and qualitative experiments demonstrate that our method outperforms mainstream baseline models across multiple metrics in the task of attributing harmful memes in Chinese. The code, Ex-ToxiCN-MM dataset, and Chinese Harmful Semantic Knowledge Base (C-HarmKB) involved in this study have been open-sourced at https://github.com/wimiw123/Ex-ToxiCN-MM

cs.CL

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single supervision signal. This paradigm overlooks the intrinsic uncertainty of low-agreement samples and often induces overconfident predictions, undermining reliability and generalization in complex subjective settings. In this work, we advocate uncertainty-aware subjectivity analysis, where models are expected to make predictions while expressing uncertainty that reflects human disagreement. To operationalize this perspective, we propose a two-phase Disagreement Perception and Uncertainty Alignment (DPUA) framework. Specifically, DPUA jointly models label prediction, rationale generation, and uncertainty expression under an uncertainty-aware setting. In the disagreement perception phase, adaptive decoupled learning enhances the model's sensitivity to disagreement-related cues while preserving task performance. In the uncertainty alignment phase, GRPO-based reward optimization further improves uncertainty-aware reasoning and aligns the model's confidence expression with the human disagreement distribution. Experiments on three subjectivity analysis tasks show that DPUA preserves task performance while better aligning model uncertainty with human disagreement, mitigating overconfidence on boundary samples, and improving out-of-distribution generalization.

cs.CL

MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding

MER2026 marks the fourth edition of the MER series of challenges. The MER series provides valuable data resources to the research community and offers tasks centered on recent research trends, establishing itself as one of the largest challenges in the field. Throughout its history, the focus of MER has shifted from discriminative emotion recognition to generative emotion understanding. Specifically, MER2023 concentrated on discriminative emotion recognition, restricting the emotion recognition scope to fixed basic labels. In MER2024 and MER2025, we transitioned to generative emotion understanding and introduced two new tasks: fine-grained emotion recognition and descriptive emotion analysis, aiming to leverage the extensive vocabulary and multimodal understanding capabilities of Multimodal Large Language Models (MLLMs) to facilitate fine-grained and explainable emotion recognition. Building on this trajectory, MER2026 continues to follow these research trends and contains four tracks: MER-Cross shifts the focus from individual to dyadic interaction scenarios; MER-FG centers on fine-grained emotion recognition; MER-Prefer aims to predict human preferences regarding different emotion descriptions; MER-PS focuses on emotion recognition based on physiological signals. More details regarding the dataset and baselines are available at https://zeroqiaoba.github.io/MER-Challenge/.

cs.HC

From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage

DNA-based storage has emerged as a promising approach to the global data crisis, offering molecular-scale density and millennial-scale stability at low maintenance cost. Over the past decade, substantial progress has been made in storing text, images, and files in DNA -- yet video remains an open challenge. The difficulty is not merely technical: effective video DNA storage requires co-designing compression and molecular encoding from the ground up, a challenge that sits at the intersection of two fields that have largely evolved independently. In this work, we present HELIX, the first end-to-end neural network jointly optimizing video compression and DNA encoding -- prior approaches treat the two stages independently, leaving biochemical constraints and compression objectives fundamentally misaligned. Our key insight: token-based representations naturally align with DNA's quaternary alphabet -- discrete semantic units map directly to ATCG bases. We introduce TK-SCONE (Token-Kronecker Structured Constraint-Optimized Neural Encoding), which achieves 1.91 bits per nucleotide through Kronecker-structured mixing that breaks spatial correlations and FSM-based mapping that guarantees biochemical constraints. Unlike two-stage approaches, HELIX learns token distributions simultaneously optimized for visual quality, prediction under masking, and DNA synthesis efficiency. This work demonstrates for the first time that learned compression and molecular storage converge naturally at token representations -- suggesting a new paradigm where neural video codecs are designed for biological substrates from the ground up.

cs.CV

Propagating Similarity, Mitigating Uncertainty: Similarity Propagation-enhanced Uncertainty for Multimodal Recommendation

Multimodal Recommendation (MMR) systems are crucial for modern platforms but are often hampered by inherent noise and uncertainty in modal features, such as blurry images, diverse visual appearances, or ambiguous text. Existing methods often overlook this modality-specific uncertainty, leading to ineffective feature fusion. Furthermore, they fail to leverage rich similarity patterns among users and items to refine representations and their corresponding uncertainty estimates. To address these challenges, we propose a novel framework, Similarity Propagation-enhanced Uncertainty for Multimodal Recommendation (SPUMR). SPUMR explicitly models and mitigates uncertainty by first constructing the Modality Similarity Graph and the Collaborative Similarity Graph to refine representations from both content and behavioral perspectives. The Uncertainty-aware Preference Aggregation module then adaptively fuses the refined multimodal features, assigning greater weight to more reliable modalities. Extensive experiments on three benchmark datasets demonstrate that SPUMR achieves significant improvements over existing leading methods.

cs.IR