SearcharxivSearch

arXiv subjects

Yu Ran

Publications and source records attributed to Yu Ran.

8 recordsLinked to original sources

MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models

Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coordinate generation process have been largely overlooked. We observe that each coordinate digit is predicted as a categorical token, yet after parsing, changing a hundreds-place digit by one changes the corresponding numerical coordinate component by 100 units, which can induce a large displacement of the executed click. This observation motivates attack objectives that account for the numerical and place-value structure of coordinate outputs rather than treating them as ordinary text. Moreover, untargeted and targeted attacks impose different success conditions--displacing the click outside the correct region versus into an attacker-specified region--and therefore benefit from different objectives. We propose MissClick, a simple and effective white-box adversarial attack with two goal-specific objectives: MissClick-U maximizes soft-coordinate displacement for untargeted disruption, while MissClick-T minimizes a place-weighted target-digit loss for targeted hijacking. Compared with existing attacks against GUI grounding models on OS-Atlas and UGround across desktop, web, and mobile platforms, MissClick-U achieves untargeted success rates of 75.07\% and 72.93\% (+16.62 and +30.72 pp), and MissClick-T achieves targeted success rates of 44.86\% and 62.67\% (+31.73 and +47.06 pp). Attack objective comparison further shows that soft-coordinate displacement yields the highest untargeted attack success rate, whereas place-weighted target-digit optimization yields the highest targeted attack success rate, revealing distinct objective preferences for the two attack goals.

cs.AI

Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality

Hallucination in large language models (LLMs) during long-form generation remains difficult to address under existing reinforcement learning from human feedback (RLHF) frameworks, as their preference rewards often overlook the model's own knowledge boundaries. In this paper, we propose the $\textbf{K}$nowledge-$\textbf{L}$evel $\textbf{C}$onsistency Reinforcement Learning $\textbf{F}$ramework ($\textbf{KLCF}$), which re-examines this problem from a distribution alignment perspective. KLCF formalizes long-form factuality as a bidirectional distribution matching objective between the policy model's expressed knowledge distribution and the base model's parametric knowledge distribution: under the constraint that generation must not exceed the support set of the base knowledge, the objective maximizes coverage of high-probability facts, thereby jointly optimizing precision and recall. To achieve this, we design a Dual-Fact Alignment mechanism that approximates the recall term using a factual checklist constructed by sampling from the base model, and constrains hallucinations with a lightweight truthfulness reward model. Both components are jointly optimized and require no external retrieval throughout training. Experimental results demonstrate that KLCF consistently improves factuality metrics across multiple long-form benchmarks and model scales, effectively alleviating hallucination and over-conservatism while maintaining efficiency and scalability.

cs.CL

Secure Video Quality Assessment Resisting Adversarial Attacks

The exponential surge in video traffic has intensified the imperative for Video Quality Assessment (VQA). Leveraging cutting-edge architectures, current VQA models have achieved human-comparable accuracy. However, recent studies have revealed the vulnerability of existing VQA models against adversarial attacks. To establish a reliable and practical assessment system, a secure VQA model capable of resisting such malicious attacks is urgently demanded. Unfortunately, no attempt has been made to explore this issue. This paper first attempts to investigate general adversarial defense principles, aiming at endowing existing VQA models with security. Specifically, we first introduce random spatial grid sampling on the video frame for intra-frame defense. Then, we design pixel-wise randomization through a guardian map, globally neutralizing adversarial perturbations. Meanwhile, we extract temporal information from the video sequence as compensation for inter-frame defense. Building upon these principles, we present a novel VQA framework from the security-oriented perspective, termed SecureVQA. Extensive experiments indicate that SecureVQA sets a new benchmark in security while achieving competitive VQA performance compared with state-of-the-art models. Ablation studies delve deeper into analyzing the principles of SecureVQA, demonstrating their generalization and contributions to the security of leading VQA models.

cs.CV

Black-box Adversarial Attacks Against Image Quality Assessment Models

The goal of No-Reference Image Quality Assessment (NR-IQA) is to predict the perceptual quality of an image in line with its subjective evaluation. To put the NR-IQA models into practice, it is essential to study their potential loopholes for model refinement. This paper makes the first attempt to explore the black-box adversarial attacks on NR-IQA models. Specifically, we first formulate the attack problem as maximizing the deviation between the estimated quality scores of original and perturbed images, while restricting the perturbed image distortions for visual quality preservation. Under such formulation, we then design a Bi-directional loss function to mislead the estimated quality scores of adversarial examples towards an opposite direction with maximum deviation. On this basis, we finally develop an efficient and effective black-box attack method against NR-IQA models. Extensive experiments reveal that all the evaluated NR-IQA models are vulnerable to the proposed attack method. And the generated perturbations are not transferable, enabling them to serve the investigation of specialities of disparate IQA models.

cs.CV

Vulnerabilities in Video Quality Assessment Models: The Challenge of Adversarial Attacks

No-Reference Video Quality Assessment (NR-VQA) plays an essential role in improving the viewing experience of end-users. Driven by deep learning, recent NR-VQA models based on Convolutional Neural Networks (CNNs) and Transformers have achieved outstanding performance. To build a reliable and practical assessment system, it is of great necessity to evaluate their robustness. However, such issue has received little attention in the academic community. In this paper, we make the first attempt to evaluate the robustness of NR-VQA models against adversarial attacks, and propose a patch-based random search method for black-box attack. Specifically, considering both the attack effect on quality score and the visual quality of adversarial video, the attack problem is formulated as misleading the estimated quality score under the constraint of just-noticeable difference (JND). Built upon such formulation, a novel loss function called Score-Reversed Boundary Loss is designed to push the adversarial video's estimated quality score far away from its ground-truth score towards a specific boundary, and the JND constraint is modeled as a strict $L_2$ and $L_\infty$ norm restriction. By this means, both white-box and black-box attacks can be launched in an effective and imperceptible manner. The source code is available at https://github.com/GZHU-DVL/AttackVQA.

cs.CV

JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs

Existing pre-trained models for knowledge-graph-to-text (KG-to-text) generation simply fine-tune text-to-text pre-trained models such as BART or T5 on KG-to-text datasets, which largely ignore the graph structure during encoding and lack elaborate pre-training tasks to explicitly model graph-text alignments. To tackle these problems, we propose a graph-text joint representation learning model called JointGT. During encoding, we devise a structure-aware semantic aggregation module which is plugged into each Transformer layer to preserve the graph structure. Furthermore, we propose three new pre-training tasks to explicitly enhance the graph-text alignment including respective text / graph reconstruction, and graph-text alignment in the embedding space via Optimal Transport. Experiments show that JointGT obtains new state-of-the-art performance on various KG-to-text datasets.

cs.CL

Non-homogeneous Problems for Nonlinear Schr\"odinger Equations in a Strip Domain

This paper studies the initial-boundary-value problem (IBVP) of a nonlinear Schr\"odinger equation posed on a strip domain $\mathbb{R}\times[0,1]$ with non-homogeneous Dirichlet boundary conditions. For any $s\ge0$, if the initial data $\varphi(x,y)$ is in Sobolev space $H^s(\mathbb{R}\times[0,1])$ and the boundary data $h(x,t)$ is in $$ {\cal H}^s (\mathbb{R} ) = \left \{ h (x, t) \in L^2 ( \mathbb{R}^2 ) \ \big | \ ( 1 + |\lambda | + |\xi|)^{\frac12} ( 1+ |\lambda | + |\xi |^2 )^{\frac{s}{2}}\hat h ( \lambda, \xi ) \in L^2 (\mathbb{R}^2 ) \right \} $$ where $\hat h $ is the Fourier transform of $h$ with respect to $t$ and $ x$, the local well-posedness of the IBVP in $C([0,T]; H^s(\mathbb{R} \times [0,1]))$ is proved. The global well-posedness is also obtained for $s = 1$. The basic idea used here relies on the derivation of an integral operator for the non-homogeneous boundary data and the proof of the series version of Strichartz's estimates for this operator. After the problem is transformed to finding a fixed point of an integral operator, the contraction mapping argument then yields a fixed point using the Strichartz's estimates for initial and boundary operators. The global well-posedness is proved using {\it a-priori} estimates of the solutions.

math.AP

Nonhomogeneous Boundary Value Problems of Nonlinear Schr\"odinger Equations in a Half Plane

This paper discusses the initial-boundary-value problems (IBVP) of nonlinear Schr\"odinger equations posed in a half plane $\mathbb{R} \times \mathbb{R}^+$ with nonhomogeneous Dirichlet boundary conditions. For any given $s \ge 0$, if the initial data $\varphi (x, y)$ are in Sobolev space $H^s(\mathbb{R}\times \mathbb{R}^+) $ with the boundary data $ h ( x, t) $ in an optimal space ${\cal H}^s(0,T)$ as defined in the introduction, which is slightly weaker than the space $$H^{(2s+1)/4}_{t} ([0, T]; L_x^2(\mathbb{R} ) ) \cap L^2_t ( [ 0, T]; H^{s+ 1/2} _x ( \mathbb{R} ) ),$$ the local well-posedness of the IBVP in $ C ( [0, T] ; H^s ( \mathbb{R}\times \mathbb{R}^+ ) )$ is proved. The global well-posedness is also discussed for $s = 1$. The main idea of the proof is to derive a boundary integral operator for the corresponding nonhomogeneous boundary condition and obtain the Strichartz's estimates for this operator. The results presented in the paper hold for the IBVP posed in a half space $ \mathbb{R}^n\times \mathbb{R}^+$ with any $n>1$.

math.AP