SearcharxivSearch

arXiv subjects

Jianquan Yang

Publications and source records attributed to Jianquan Yang.

7 recordsLinked to original sources

SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

Text-driven image editing has advanced rapidly, but reliably localizing these manipulations requires image manipulation localization (IML) models trained on large pixel-annotated datasets, and there is still no low-cost way to obtain such training data at scale. We observe that these data already exist in disguise: public editing datasets contain millions of structurally identical (original, edited) pairs to IML training samples, lacking only pixel-level masks. Recovering these masks automatically is non-trivial: pixel differencing is overwhelmed by diffusion-induced perturbations across all pixels, and instruction-only grounding localizes only what the prompt describes, missing unintended editor side-effects. We propose SIGMA (Semantic-difference Instruction-Grounding Mask Annotator), which performs semantic-feature differencing in a vision foundation backbone and injects an instruction-derived spatial prior into this visual stream via bidirectional cross-modal refinement, amplifying the difference signal at intended-edit regions when the editor faithfully realizes user intent. SIGMA is trained in two complementary stages: Stage I supervises on inpainting masks; Stage II closes the diffusion-domain shift via VAE-roundtrip noise calibration, EMA self-training, and an edit-noise disentanglement loss. SIGMA outperforms existing automatic mask generators on five benchmarks (+12.20% F1, +11.16% IoU). When applied to public editing corpora, it produces a ~1.1M IML training set that improves six diverse detectors by +18.34% F1 across five datasets, turning previously unused editing data into a model-agnostic supervisory resource for IML. We'll release the full codebase as soon as the paper is accepted.

cs.CV

Unveiling the Attribute Misbinding Threat in Identity-Preserving Models

Identity-preserving models have led to notable progress in generating personalized content. Unfortunately, such models also exacerbate risks when misused, for instance, by generating threatening content targeting specific individuals. This paper introduces the \textbf{Attribute Misbinding Attack}, a novel method that poses a threat to identity-preserving models by inducing them to produce Not-Safe-For-Work (NSFW) content. The attack's core idea involves crafting benign-looking textual prompts to circumvent text-filter safeguards and leverage a key model vulnerability: flawed attribute binding that stems from its internal attention bias. This results in misattributing harmful descriptions to a target identity and generating NSFW outputs. To facilitate the study of this attack, we present the \textbf{Misbinding Prompt} evaluation set, which examines the content generation risks of current state-of-the-art identity-preserving models across four risk dimensions: pornography, violence, discrimination, and illegality. Additionally, we introduce the \textbf{Attribute Binding Safety Score (ABSS)}, a metric for concurrently assessing both content fidelity and safety compliance. Experimental results show that our Misbinding Prompt evaluation set achieves a \textbf{5.28}\% higher success rate in bypassing five leading text filters (including GPT-4o) compared to existing main-stream evaluation sets, while also demonstrating the highest proportion of NSFW content generation. The proposed ABSS metric enables a more comprehensive evaluation of identity-preserving models by concurrently assessing both content fidelity and safety compliance.

cs.CR

MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training

Accurately assessing the perceptual quality of face images is crucial, especially with the rapid progress in face restoration and generation. Traditional quality assessment methods often struggle with the unique characteristics of face images, limiting their generalizability. While learning-based approaches demonstrate superior performance due to their strong fitting capabilities, their high complexity typically incurs significant computational and storage costs, hindering practical deployment. To address this, we propose a lightweight face quality assessment network with Multi-Stage Progressive Training (MSPT). Our network employs a three-stage progressive training strategy that gradually introduces more diverse data samples and increases input image resolution. This novel approach enables lightweight networks to achieve high performance by effectively learning complex quality features while significantly mitigating catastrophic forgetting. Our MSPT achieved the second highest score on the VQualA 2025 face image quality assessment benchmark dataset, demonstrating that MSPT achieves comparable or better performance than state-of-the-art methods while maintaining efficient inference.

cs.MM

Decay estimates for the compressible viscoelastic equations in an exterior domain

In this paper, we study the compressible viscoelastic equations in an exterior domain. We prove the $L_2$ estimates for the solution to the linearized problem and show the decay estimates for the solution to the nonlinear problem. In particular, we obtain the optimal decay rates of the solution itself and its spatial-time derivatives in the $L_2$-norm.

math.AP

NTIRE 2024 Quality Assessment of AI-Generated Content Challenge

This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2024. This challenge is to address a major challenge in the field of image and video processing, namely, Image Quality Assessment (IQA) and Video Quality Assessment (VQA) for AI-Generated Content (AIGC). The challenge is divided into the image track and the video track. The image track uses the AIGIQA-20K, which contains 20,000 AI-Generated Images (AIGIs) generated by 15 popular generative models. The image track has a total of 318 registered participants. A total of 1,646 submissions are received in the development phase, and 221 submissions are received in the test phase. Finally, 16 participating teams submitted their models and fact sheets. The video track uses the T2VQA-DB, which contains 10,000 AI-Generated Videos (AIGVs) generated by 9 popular Text-to-Video (T2V) models. A total of 196 participants have registered in the video track. A total of 991 submissions are received in the development phase, and 185 submissions are received in the test phase. Finally, 12 participating teams submitted their models and fact sheets. Some methods have achieved better results than baseline methods, and the winning methods in both tracks have demonstrated superior prediction performance on AIGC.

cs.CV

Collimation method studies for next-generation hadron colliders

In order to handle extremely-high stored energy in future proton-proton colliders, an extremely high-efficiency collimation system is required for safe operation. At LHC, the major limiting locations in terms of particle losses on superconducting (SC) magnets are the dispersion suppressors (DS) downstream of the transverse collimation insertion. These losses are due to the protons experiencing single diffractive interactions in the primary collimators. How to solve this problem is very important for future proton-proton colliders, such as the FCC-hh and SPPC. In this article, a novel method is proposed, which arranges both the transverse and momentum collimation in the same long straight section. In this way, the momentum collimation system can clean those particles related to the single diffractive effect. The effectiveness of the method has been confirmed by multi-particle simulations. In addition, SC quadrupoles with special designs such as enlarged aperture and good shielding are adopted to enhance the phase advance in the transverse collimation section, so that tertiary collimators can be arranged to clean off the tertiary halo which emerges from the secondary collimators and improve the collimation efficiency. With one more collimation stage in the transverse collimation, the beam losses in both the momentum collimation section and the experimental regions can be largely reduced. Multi-particle simulation results with the MERLIN code confirm the effectiveness of the collimation method. At last, we provide a protection scheme of the SC magnets in the collimation section. The FLUKA simulations show that by adding some special protective collimators in front of the magnets, the maximum power deposition in the SC coils is reduced dramatically, which is proven to be valid for protecting the SC magnets from quenching.

physics.acc-ph

Resonant slow extraction in synchrotrons by using anti-symmetric sextupole fields

This paper proposes a novel method for resonant slow extraction in synchrotrons by using special anti-symmetric sextupole fields, which can be produced by a special magnet structure. The method has the potential in applications demanding for very stable slow extraction from synchrotrons. Our studies show that the slow extraction at the half-integer resonance by using anti-symmetric sextupole field has some advantages compared to the normal sextupole field, and the latter is widely used in the slow extraction method. One of them is that it can work at a more distant tune from the resonance, so that it can reduce significantly the intensity variation of the extracted beam which is mainly caused by the ripples of magnet power supplies. The studies by both the Hamiltonian theory and numerical simulations show that the stable region at the proximity of the half-integer resonance by anti-symmetric sextupole field is much smaller and flatter than the one by standard sextupole field at the third-order resonance. By gradually increasing the field strength, the beam can be extracted with intensity more homogeneous than by the usual third-order resonant method, in the means of both smaller intensity variation and spike in the beginning spill. Similar to the case with a normal sextupole, we derive an empirical formula for the area of the stable region with an anti-symmetric sextupole. One can find that with the same field strength and the same tune distance to the resonance, the area of stable region or the change of the area due to the working point variation in the case of anti-symmetric sextupole is about 1/14 of the one in the case of standard sextupole. The detailed studies including beam dynamic behaviors at the proximity of other resonances, influence of 2-D field error, half-integer stop-band, and resonant slow extraction by using quadrupole field have also been presented.

physics.acc-ph