SearcharxivSearch

arXiv subjects

Zhixin Xie

Publications and source records attributed to Zhixin Xie.

At least 19 recordsLinked to original sources

Generalised Cone Conjecture, II: Generic Sakai surfaces

In this paper we prove that the Generalised Cone Conjecture holds for generic Sakai surfaces, which completes the proof of the Generalised Cone Conjecture in dimension two. Furthermore, we show that there exists a non-generic Sakai surface for which no meaningful variant of the conjecture holds.

math.AG

MMP for Enriques pairs and singular Enriques varieties

We introduce and study the class of primitive Enriques varieties, whose smooth members are Enriques manifolds. We provide several examples and we demonstrate that this class is stable under the operations of the Minimal Model Program (MMP). In particular, given an Enriques manifold $Y$ and an effective $\mathbb{R}$-divisor $B_Y$ on $Y$ such that the pair $(Y,B_Y)$ is log canonical, we prove that any $(K_Y+B_Y)$-MMP terminates with a minimal model $(Y',B_{Y'})$ of $(Y,B_Y)$, where $Y'$ is a $\mathbb{Q}$-factorial primitive Enriques variety with canonical singularities. Finally, we investigate the asymptotic theory of Enriques manifolds.

math.AG

TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning

The demand of customized large language models (LLMs) has led to commercial LLMs offering black-box fine-tuning APIs, yet this convenience introduces a critical security loophole: attackers could jailbreak the LLMs by fine-tuning them with malicious data. Though this security issue has recently been exposed, the feasibility of such attacks is questionable as malicious training dataset is believed to be detectable by moderation models such as Llama-Guard-3. In this paper, we propose TrojanPraise, a novel finetuning-based attack exploiting benign and thus filter-approved data. Basically, TrojanPraise fine-tunes the model to associate a crafted word (e.g., "bruaf") with harmless connotations, then uses this word to praise harmful concepts, subtly shifting the LLM from refusal to compliance. To explain the attack, we decouple the LLM's internal representation of a query into two dimensions of knowledge and attitude. We demonstrate that successful jailbreak requires shifting the attitude while avoiding knowledge shift, a distortion in the model's understanding of the concept. To validate this attack, we conduct experiments on five opensource LLMs and two commercial LLMs under strict black-box settings. Results show that TrojanPraise achieves a maximum attack success rate of 95.88% while evading moderation.

cs.CR

Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position

Diffusion Large Language Models (dLLMs) have recently emerged as a competitive non-autoregressive paradigm due to their unique training and inference approach. However, there is currently a lack of safety study on this novel architecture. In this paper, we present the first analysis of dLLMs' safety performance and propose a novel safety alignment method tailored to their unique generation characteristics. Specifically, we identify a critical asymmetry between the defender and attacker in terms of security. For the defender, we reveal that the middle tokens of the response, rather than the initial ones, are more critical to the overall safety of dLLM outputs; this seems to suggest that aligning middle tokens can be more beneficial to the defender. The attacker, on the contrary, may have limited power to manipulate middle tokens, as we find dLLMs have a strong tendency towards a sequential generation order in practice, forcing the attack to meet this distribution and diverting it from influencing the critical middle tokens. Building on this asymmetry, we introduce Middle-tOken Safety Alignment (MOSA), a novel method that directly aligns the model's middle generation with safe refusals exploiting reinforcement learning. We implement MOSA and compare its security performance against eight attack methods on two benchmarks. We also test the utility of MOSA-aligned dLLM on coding, math, and general reasoning. The results strongly prove the superiority of MOSA.

cs.CR

Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs

Despite substantial efforts in safety alignment, recent research indicates that Large Language Models (LLMs) remain highly susceptible to jailbreak attacks. Among these attacks, finetuning-based ones that compromise LLMs' safety alignment via fine-tuning stand out due to its stable jailbreak performance. In particular, a recent study indicates that fine-tuning with as few as 10 harmful question-answer (QA) pairs can lead to successful jailbreaking across various harmful questions. However, such malicious fine-tuning attacks are readily detectable and hence thwarted by moderation models. In this paper, we demonstrate that LLMs can be jailbroken by fine-tuning with only 10 benign QA pairs; our attack exploits the increased sensitivity of LLMs to fine-tuning data after being overfitted. Specifically, our fine-tuning process starts with overfitting an LLM via fine-tuning with benign QA pairs involving identical refusal answers. Further fine-tuning is then performed with standard benign answers, causing the overfitted LLM to forget the refusal attitude and thus provide compliant answers regardless of the harmfulness of a question. We implement our attack on the ten LLMs and compare it with five existing baselines. Experiments demonstrate that our method achieves significant advantages in both attack effectiveness and attack stealth. Our findings expose previously unreported security vulnerabilities in current LLMs and provide a new perspective on understanding how LLMs' security is compromised, even with benign fine-tuning. Our code is available at https://github.com/ZHIXINXIE/tenBenign.

cs.CR

Nakayama-Zariski decomposition and the termination of flips

We show that for pseudoeffective projective pairs the termination of one sequence of flips implies the termination of all flips, assuming a natural conjecture on the behaviour of the Nakayama-Zariski decomposition under the operations of a Minimal Model Program.

math.AG

Dagger Behind Smile: Fool LLMs with a Happy Ending Story

The wide adoption of Large Language Models (LLMs) has attracted significant attention from $\textit{jailbreak}$ attacks, where adversarial prompts crafted through optimization or manual design exploit LLMs to generate malicious contents. However, optimization-based attacks have limited efficiency and transferability, while existing manual designs are either easily detectable or demand intricate interactions with LLMs. In this paper, we first point out a novel perspective for jailbreak attacks: LLMs are more responsive to $\textit{positive}$ prompts. Based on this, we deploy Happy Ending Attack (HEA) to wrap up a malicious request in a scenario template involving a positive prompt formed mainly via a $\textit{happy ending}$, it thus fools LLMs into jailbreaking either immediately or at a follow-up malicious request. This has made HEA both efficient and effective, as it requires only up to two turns to fully jailbreak LLMs. Extensive experiments show that our HEA can successfully jailbreak on state-of-the-art LLMs, including GPT-4o, Llama3-70b, Gemini-pro, and achieves 88.79% attack success rate on average. We also provide quantitative explanations for the success of HEA.

cs.CL

Deformations of primitive Enriques varieties

We develop the deformation theory of primitive Enriques varieties, which are defined as quasi-étale quotients of primitive symplectic varieties by nonsymplectic group actions. In particular, we establish a local Torelli theorem for primitive Enriques varieties. As applications thereof, we describe the behavior of certain primitive Enriques varieties under locally trivial deformations.

math.AG

On the relative cone conjecture for families of IHS manifolds

We study the relative cone conjecture for families of $K$-trivial varieties with vanishing irregularity. As an application we prove that the relative movable and the relative nef cone conjectures hold for fibrations in projective IHS manifolds of the 4 known deformation types.

math.AG

Shaking the Fake: Detecting Deepfake Videos in Real Time via Active Probes

Real-time deepfake, a type of generative AI, is capable of "creating" non-existing contents (e.g., swapping one's face with another) in a video. It has been, very unfortunately, misused to produce deepfake videos (during web conferences, video calls, and identity authentication) for malicious purposes, including financial scams and political misinformation. Deepfake detection, as the countermeasure against deepfake, has attracted considerable attention from the academic community, yet existing works typically rely on learning passive features that may perform poorly beyond seen datasets. In this paper, we propose SFake, a new real-time deepfake detection method that innovatively exploits deepfake models' inability to adapt to physical interference. Specifically, SFake actively sends probes to trigger mechanical vibrations on the smartphone, resulting in the controllable feature on the footage. Consequently, SFake determines whether the face is swapped by deepfake based on the consistency of the facial area with the probe pattern. We implement SFake, evaluate its effectiveness on a self-built dataset, and compare it with six other detection methods. The results show that SFake outperforms other detection methods with higher detection accuracy, faster process speed, and lower memory consumption.

cs.CV

Parallel Ising Annealer via Gradient-based Hamiltonian Monte Carlo

Ising annealer is a promising quantum-inspired computing architecture for combinatorial optimization problems. In this paper, we introduce an Ising annealer based on the Hamiltonian Monte Carlo, which updates the variables of all dimensions in parallel. The main innovation is the fusion of an approximate gradient-based approach into the Ising annealer which introduces significant acceleration and allows a portable and scalable implementation on the commercial FPGA. Comprehensive simulation and hardware experiments show that the proposed Ising annealer has promising performance and scalability on all types of benchmark problems when compared to other Ising annealers including the state-of-the-art hardware. In particular, we have built a prototype annealer which solves Ising problems of both integer and fraction coefficients with up to 200 spins on a single low-cost FPGA board, whose performance is demonstrated to be better than the state-of-the-art quantum hardware D-Wave 2000Q and similar to the expensive coherent Ising machine. The sub-linear scalability of the annealer signifies its potential in solving challenging combinatorial optimization problems and evaluating the advantage of quantum hardware.

quant-ph

Comparison and uniruledness of asymptotic base loci

We prove that the asymptotic base loci of an NQC klt generalized pair with big canonical class are uniruled. We also show that the non-nef locus and the diminished base locus of the adjoint divisor of an NQC log canonical generalized pair coincide. As applications, we study the uniruledness of the asymptotic base loci associated with pseudo-effective divisors on generalized log Calabi-Yau type varieties.

math.AG

Rigid currents in birational geometry

A rigid current on a compact complex manifold is a closed positive current whose cohomology class contains only one closed positive current. Rigid currents occur in complex dynamics, algebraic and differential geometry. The goals of the present paper are: (a) to give a systematic treatment of rigid currents, (b) to demonstrate how they appear within the Minimal Model Program, and (c) to give many new examples of rigid currents.

math.AG