SearcharxivSearch

arXiv subjects

Xinyi Wei

Publications and source records attributed to Xinyi Wei.

7 recordsLinked to original sources

VTONQA: A Multi-Dimensional Quality Assessment Dataset for Virtual Try-on

With the rapid development of e-commerce and digital fashion, image-based virtual try-on (VTON) has attracted increasing attention. However, existing VTON models often suffer from artifacts such as garment distortion and body inconsistency, highlighting the need for reliable quality evaluation of VTON-generated images. To this end, we construct \textbf{VTONQA}, the first multi-dimensional quality assessment dataset specifically designed for VTON, which contains 8,132 images generated by 11 representative VTON models, along with 24,396 mean opinion scores (MOSs) across three evaluation dimensions (\textit{i.e.}, clothing fit, body compatibility, and overall quality). Based on VTONQA, we benchmark both VTON models and a diverse set of image quality assessment (IQA) metrics, revealing the limitations of existing methods and highlighting the value of the proposed dataset. We believe that the VTONQA dataset and corresponding benchmarks will provide a solid foundation for perceptually aligned evaluation, benefiting both the development of quality assessment methods and the advancement of VTON models. The dataset we proposed in this paper is publicly available at: https://huggingface.co/datasets/weixiny0408/vtonqa.

cs.CV

Higher-Order Boundary Conditions for Atomistic Dislocation Simulations

We present a higher-order boundary condition for atomistic simulations of dislocations that address the slow convergence of standard supercell methods. The method is based on a multipole expansion of the equilibrium displacement, combining continuum predictor solutions with discrete moment corrections. The continuum predictors are computed by solving a hierarchy of singular elliptic PDEs via a Galerkin spectral method, while moment coefficients are evaluated from force-moment identities with controlled approximation error. A key feature is the coupling between accurate continuum predictors and moment evaluations, enabling the construction of systematically improvable high-order boundary conditions. We thus design novel algorithms, and numerical results for screw and edge dislocations confirm the predicted convergence rates in geometry and energy norms, with reduced finite-size effects and moderate computational cost.

math.NA

Planning Stealthy Backdoor Attacks in MDPs with Observation-Based Triggers

This paper investigates backdoor attack planning in stochastic control systems modeled as Markov Decision Processes (MDPs). A backdoor attack involves an adversary deploying a policy that performs well in the original MDP to pass testing, but behaves maliciously at runtime when combined with a trigger that perturbs system dynamics. We consider a sophisticated attacker capable of jointly optimizing the backdoor policy and its trigger using only a blackbox simulator. During execution, the attacker has access only to partial observations of the system state and is restricted to introduce small perturbations to the system's transition dynamics. We formulate the attack planning problem as a constrained Markov game with an augmented state space and two players: Player 0 learns a backdoor policy that maximizes attack rewards when the trigger is active. However, when the trigger is inactive, the backdoor policy behaves near-optimally in the original MDP; Player 1 designs a finite-memory, observation-based trigger to activate the attack. We propose a switching gradient-based optimization algorithm to jointly solve for the backdoor policy and trigger. Experiments on a case study demonstrate the effectiveness of our method in achieving stealthy and successful backdoor attacks, and how the attack performance varies under different parameters related to the stealthiness of the backdoor attack.

eess.SY

Active Inference through Incentive Design in Markov Decision Processes

We present a method for active inference with partial observations in stochastic systems through incentive design, also known as the leader-follower game. Consider a leader agent who aims to infer a follower agent's type given a finite set of possible types. Different types of followers differ in either the dynamical model, the reward function, or both. We assume the leader can partially observe a follower's behavior in the stochastic system modeled as a Markov decision process, in which the follower takes an optimal policy to maximize a total reward. To improve inference accuracy and efficiency, the leader can offer side payments (incentives) to the followers such that different types of them, under the incentive design, can exhibit diverging behaviors that facilitate the leader's inference task. We show the problem of active inference through incentive design can be formulated as a special class of leader-follower games, where the leader's objective is to balance the information gain and cost of incentive design. The information gain is measured by the entropy of the estimated follower's type given partial observations. Furthermore, we demonstrate that this problem can be solved by reducing a single-level optimization through softmax temporal consistency between followers' policies and value functions. This reduction allows us to develop an efficient gradient-based algorithm. We utilize observable operators in the hidden Markov model (HMM) to compute the necessary gradients and demonstrate the effectiveness of our approach through experiments in stochastic grid world environments.

eess.SY

IllusionBench+: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models

Current Visual Language Models (VLMs) show impressive image understanding but struggle with visual illusions, especially in real-world scenarios. Existing benchmarks focus on classical cognitive illusions, which have been learned by state-of-the-art (SOTA) VLMs, revealing issues such as hallucinations and limited perceptual abilities. To address this gap, we introduce IllusionBench, a comprehensive visual illusion dataset that encompasses not only classic cognitive illusions but also real-world scene illusions. This dataset features 1,051 images, 5,548 question-answer pairs, and 1,051 golden text descriptions that address the presence, causes, and content of the illusions. We evaluate ten SOTA VLMs on this dataset using true-or-false, multiple-choice, and open-ended tasks. In addition to real-world illusions, we design trap illusions that resemble classical patterns but differ in reality, highlighting hallucination issues in SOTA models. The top-performing model, GPT-4o, achieves 80.59% accuracy on true-or-false tasks and 76.75% on multiple-choice questions, but still lags behind human performance. In the semantic description task, GPT-4o's hallucinations on classical illusions result in low scores for trap illusions, even falling behind some open-source models. IllusionBench is, to the best of our knowledge, the largest and most comprehensive benchmark for visual illusions in VLMs to date.

cs.CV

Amplitude Expansion Phase Field Crystal (APFC) Modeling based Efficient Dislocation Simulations using Fourier Pseudospectral Method

Crystalline defects critically influence material properties, necessitating accurate simulation methods. Existing approaches, from atomic-scale configurations to continuum elasticity, face inherent limitations in modeling dislocation-induced lattice deformation. The amplitude expansion of the phase field crystal (APFC) model bridges this gap with a mesoscopic description. This paper introduces a computationally efficient Fourier pseudospectral method for solving the APFC equations. The method exploits system periodicity and solution analyticity--the latter's rigorous proof remaining an open question, as discussed herein--to enable precise implementation of periodic boundary conditions. Numerical experiments on 2D triangular and 3D body-centered cubic lattices demonstrate that the method accurately reproduces the strain fields of edge dislocations, matching continuum theory predictions. These results confirm the APFC model's potential for capturing complex defect structures at the mesoscale, paving the way for simulating more intricate defect dynamics.

cond-mat.mtrl-sci

Deep Unfolding with Normalizing Flow Priors for Inverse Problems

Many application domains, spanning from computational photography to medical imaging, require recovery of high-fidelity images from noisy, incomplete or partial/compressed measurements. State of the art methods for solving these inverse problems combine deep learning with iterative model-based solvers, a concept known as deep algorithm unfolding. By combining a-priori knowledge of the forward measurement model with learned (proximal) mappings based on deep networks, these methods yield solutions that are both physically feasible (data-consistent) and perceptually plausible. However, current proximal mappings only implicitly learn such image priors. In this paper, we propose to make these image priors fully explicit by embedding deep generative models in the form of normalizing flows within the unfolded proximal gradient algorithm. We demonstrate that the proposed method outperforms competitive baselines on various image recovery tasks, spanning from image denoising to inpainting and deblurring.

eess.IV