SearcharxivSearch

arXiv subjects

An-an Liu

Publications and source records attributed to An-an Liu.

6 recordsLinked to original sources

Group Relative Attention Guidance for Image Editing

Recently, image editing based on Diffusion-in-Transformer models has undergone rapid development. However, existing editing methods often lack effective control over the degree of editing, limiting their ability to achieve more customized results. To address this limitation, we investigate the MM-Attention mechanism within the DiT model and observe that the Query and Key tokens share a bias vector that is only layer-dependent. We interpret this bias as representing the model's inherent editing behavior, while the delta between each token and its corresponding bias encodes the content-specific editing signals. Based on this insight, we propose Group Relative Attention Guidance, a simple yet effective method that reweights the delta values of different tokens to modulate the focus of the model on the input image relative to the editing instruction, enabling continuous and fine-grained control over editing intensity without any tuning. Extensive experiments conducted on existing image editing frameworks demonstrate that GRAG can be integrated with as few as four lines of code, consistently enhancing editing quality. Moreover, compared to the commonly used Classifier-Free Guidance, GRAG achieves smoother and more precise control over the degree of editing. Our code will be released at https://github.com/little-misfit/GRAG-Image-Editing.

cs.CV

Causal Disentanglement Hidden Markov Model for Fault Diagnosis

In modern industries, fault diagnosis has been widely applied with the goal of realizing predictive maintenance. The key issue for the fault diagnosis system is to extract representative characteristics of the fault signal and then accurately predict the fault type. In this paper, we propose a Causal Disentanglement Hidden Markov model (CDHM) to learn the causality in the bearing fault mechanism and thus, capture their characteristics to achieve a more robust representation. Specifically, we make full use of the time-series data and progressively disentangle the vibration signal into fault-relevant and fault-irrelevant factors. The ELBO is reformulated to optimize the learning of the causal disentanglement Markov model. Moreover, to expand the scope of the application, we adopt unsupervised domain adaptation to transfer the learned disentangled representations to other working environments. Experiments were conducted on the CWRU dataset and IMS dataset. Relevant results validate the superiority of the proposed method.

cs.LG

Decomposed Prototype Learning for Few-Shot Scene Graph Generation

Today's scene graph generation (SGG) models typically require abundant manual annotations to learn new predicate types. Therefore, it is difficult to apply them to real-world applications with massive uncommon predicate categories whose annotations are hard to collect. In this paper, we focus on Few-Shot SGG (FSSGG), which encourages SGG models to be able to quickly transfer previous knowledge and recognize unseen predicates well with only a few examples. However, current methods for FSSGG are hindered by the high intra-class variance of predicate categories in SGG: On one hand, each predicate category commonly has multiple semantic meanings under different contexts. On the other hand, the visual appearance of relation triplets with the same predicate differs greatly under different subject-object compositions. Such great variance of inputs makes it hard to learn generalizable representation for each predicate category with current few-shot learning (FSL) methods. However, we found that this intra-class variance of predicates is highly related to the composed subjects and objects. To model the intra-class variance of predicates with subject-object context, we propose a novel Decomposed Prototype Learning (DPL) model for FSSGG. Specifically, we first construct a decomposable prototype space to capture diverse semantics and visual patterns of subjects and objects for predicates by decomposing them into multiple prototypes. Afterwards, we integrate these prototypes with different weights to generate query-adaptive predicate representation with more reliable semantics for each query sample. We conduct extensive experiments and compare with various baseline methods to show the effectiveness of our method.

cs.CV

Intrinsic Bias Identification on Medical Image Datasets

Machine learning based medical image analysis highly depends on datasets. Biases in the dataset can be learned by the model and degrade the generalizability of the applications. There are studies on debiased models. However, scientists and practitioners are difficult to identify implicit biases in the datasets, which causes lack of reliable unbias test datasets to valid models. To tackle this issue, we first define the data intrinsic bias attribute, and then propose a novel bias identification framework for medical image datasets. The framework contains two major components, KlotskiNet and Bias Discriminant Direction Analysis(bdda), where KlostkiNet is to build the mapping which makes backgrounds to distinguish positive and negative samples and bdda provides a theoretical solution on determining bias attributes. Experimental results on three datasets show the effectiveness of the bias attributes discovered by the framework.

cs.CV

Treatment of Linear and Nonlinear Dielectric Property of Molecular Monolayer and Submonolayer with Microscopic Dipole Lattice Model: I. Second Harmonic Generation and Sum-Frequency Generation

In the currently accepted models of the nonlinear optics, the nonlinear radiation was treated as the result of an infinitesimally thin polarization sheet layer, and a three layer model was generally employed. The direct consequence of this approach is that an apriori dielectric constant, which still does not have a clear definition, has to be assigned to this polarization layer. Because the Second Harmonic Generation (SHG) and the Sum-Frequency Generation vibrational Spectroscopy (SFG-VS) have been proven as the sensitive probes for interfaces with the submonolayer coverage, the treatment based on the more realistic discrete induced dipole model needs to be developed. Here we show that following the molecular optics theory approach the SHG, as well as the SFG-VS, radiation from the monolayer or submonolayer at an interface can be rigorously treated as the radiation from an induced dipole lattice at the interface. In this approach, the introduction of the polarization sheet is no longer necessary. Therefore, the ambiguity of the unaccounted dielectric constant of the polarization layer is no longer an issue. Moreover, the anisotropic two dimensional microscopic local field factors can be explicitly expressed with the linear polarizability tensors of the interfacial molecules. Based on the planewise dipole sum rule in the molecular monolayer, crucial experimental tests of this microscopic treatment with SHG and SFG-VS are discussed. Many puzzles in the literature of surface SHG and SFG spectroscopy studies can also be understood or resolved in this framework. This new treatment may provide a solid basis for the quantitative analysis in the surface SHG and SFG studies.

physics.chem-ph

A New Type of Two-photon Forward Radiation in Pure Liquids

Unexpected spectral features are observed in the two photon spectrum of the pure water in the forward direction when an 80 femtosecond laser pulse is focused at 10^10Wcm-2 or less. Such intensity is much lower than the breakdown or stimulated threshold of the liquid water. The two broad features are about 2700cm-1 and 5000cm-1 red shifted from the hyper-Rayleigh wavelength, respectively, and they are quadratic with the laser intensity. They do not match the known Raman or hyper-Raman frequencies of water, and they are both centered at a narrow angle in the forward direction. Several other liquids also exhibited similar but molecular specific spectral features.

physics.chem-ph