Searcharxiv⌕ Search

arXiv subjects

Yuxia Qiao

Publications and source records attributed to Yuxia Qiao.

3 recordsLinked to original sources

C2P-VAR: Continual and Compositional Personalization in Visual Autoregressive Models

Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation, yet their personalization capabilities remain largely limited to static, single-concept settings. In practice, users may continuously introduce new concepts and wish to compose multiple personalized concepts within a single image. Such scenarios pose two fundamental challenges: catastrophic forgetting during sequential personalization and feature interference during multi-concept composition. In this work, we study continual and compositional personalization in VAR models and propose C2P-VAR, a unified framework addressing both challenges. For continual personalization, we introduce C2PVAR-S, which identifies concept-relevant parameters from gradient magnitudes and dynamically updates their selection during training. To preserve previously learned concepts, C2PVAR-S applies regularization only to parameters shared by the current and historical concepts, thereby reducing unnecessary interference without introducing additional model components. For multi-concept personalization, we further propose C2PVAR-M, which employs parallel global and concept-specific branches with spatially localized feature fusion and logit aggregation to achieve controllable concept placement and reduce feature entanglement. Extensive experiments on continual and multi-concept personalization demonstrate that C2P-VAR consistently outperforms existing baselines in subject fidelity, while maintaining competitive text alignment and introducing negligible storage and inference overhead. Our results establish a unified framework for scalable and controllable personalization of visual autoregressive models.

cs.CV↗

HyperErase: Scale-Calibrated Hypernetwork for Multi-Concept Erasure in Text-to-Image Models

Recent advances in text-to-image (T2I) generation have substantially improved visual synthesis, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods predominantly follow a static weight paradigm, producing a single frozen adapter that struggles to adapt to diverse prompt variations and suffers from parameter interference when scaling to multiple concepts. We propose \textbf{HyperErase}, a framework for concept erasure based on hypernetwork-driven prompt-conditioned parameter synthesis. Our approach first reframes concept erasure as prompt-conditioned parameter amortization and trains a hypernetwork to map textual descriptions to prompt-specific LoRA updates, eliminating the need for per-prompt gradient optimization or manual LoRA merging. To further improve the stability and precision of synthesized adapters, we develop a decoupled rectification strategy, which disentangles LoRA tokens into pattern and scale subspaces, applies a square-root transform to curb multiplicative over-scaling, and leverages teacher-derived canonical priors for inference-time correction. Extensive experiments across major concept categories demonstrate that HyperErase consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving performance comparable to gold-standard single-concept baselines. Furthermore, the resulting models can provide specialized LoRAs for each input prompt variation in a single forward pass without requiring gradient updates during inference. These principled and flexible framework offers a new paradigm for concept erasure in T2I models.

cs.CV↗

LEMUR: Latent Entropy-aware Multimodal Unlearning via Visual-anchored Reasoning Redirection

Reinforcement-learning (RL) post-training equips multimodal large reasoning models (MLRMs) with exploratory chains of thought (CoT), substantially improving visual reasoning. However, we find that this capability introduces a distinct privacy vulnerability: even when a sensitive fact is successfully unlearned from the final answer, the model may still reproduce it in its reasoning trace. This leakage is substantially more pronounced in natively RL-trained MLRMs than in their non -reasoning base models, revealing a privacy risk that existing unlearning methods are not designed to address. We show that RL-induced exploration leaves sensitive content with a distinctive token-level entropy signature that is largely absent from base models. Based on this observation, we propose LEMUR, a fully training-free, inference-time unlearning framework for natively RL-trained multimodal models. LEMUR uses entropy dynamics as a control signal to identify when sensitive reasoning begins and when sanitization should stop. During this interval, it redirects the reasoning trajectory through entropy-modulated visual-anchor latent injection, replacing committed tokens with sanitized, probability-weighted embeddings re-grounded in the input image. Across diverse MLRMs, LEMUR consistently outperforms existing unlearning met hods in suppressing both reasoning-trace and answer leakage, while better preserving non-sensitive utility and output fluency. These results demonstrate that RL-induced entropy dynamics provide a distinctive signal for privacy leakage and that exploiting this signal enables effective training-free unlearning for reasoning-capable multimodal models.

cs.LG↗