SearcharxivSearch

arXiv subjects

Sheng-Yu Wang

Publications and source records attributed to Sheng-Yu Wang.

18 recordsLinked to original sources

Core-Level Spectroscopy Decodes Bond-Alternation Dynamics of Cyclo[18]Carbon

The advent of X-ray free-electron lasers and high-harmonic generation has made time-resolved X-ray spectroscopy a powerful tool for probing local atomic environments, yet whether localized core excitations can report on global collective distortions remains open. Cyclo[18]carbon (C$_{18}$), with its polyynic ground state (D$_\text{9h}$) and cumulenic transition state (D$_\text{18h}$), provides an ideal model to address this long-standing issue in bond-length alternation (BLA) dynamics. Mapping two-dimensional potential energy surfaces by first-principles simulations, we find that core ionization symmetrizes the ground-state double-well potential along the BLA coordinate. Our calculated X-ray spectra reveal remarkable sensitivity to bond-length variations: C1s ionization potentials vary by up to 2.4~eV across the BLA coordinate (1.1--1.4~\AA), with a 0.9~eV variation for minima predicted by different functionals, while NEXAFS $\pi^*$ peaks shift by up to 4~eV across the same coordinate. These predicted signatures provide a quantitative spectroscopy--structure dictionary for decoding transient structures in future ultrafast X-ray experiments and monitoring bond-alternation dynamics in real time.

physics.atm-clus

The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric

Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS

cs.CV

Fast Data Attribution for Text-to-Image Models

Data attribution for text-to-image models aims to identify the training images that most significantly influenced a generated output. Existing attribution methods involve considerable computational resources for each query, making them impractical for real-world applications. We propose a novel approach for scalable and efficient data attribution. Our key idea is to distill a slow, unlearning-based attribution method to a feature embedding space for efficient retrieval of highly influential training images. During deployment, combined with efficient indexing and search methods, our method successfully finds highly influential images without running expensive attribution algorithms. We show extensive results on both medium-scale models trained on MSCOCO and large-scale Stable Diffusion models trained on LAION, demonstrating that our method can achieve better or competitive performance in a few seconds, faster than existing methods by 2,500x - 400,000x. Our work represents a meaningful step towards the large-scale application of data attribution methods on real-world models such as Stable Diffusion.

cs.CV

Learning an Image Editing Model without Image Editing Pairs

Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Current workarounds use synthetic training pairs that leverage the zero-shot capabilities of existing models. However, this can propagate and magnify the artifacts of the pretrained model into the final trained model. In this work, we present a new training paradigm that eliminates the need for paired data entirely. Our approach directly optimizes a few-step diffusion model by unrolling it during training and leveraging feedback from vision-language models (VLMs). For each input and editing instruction, the VLM evaluates if an edit follows the instruction and preserves unchanged content, providing direct gradients for end-to-end optimization. To ensure visual fidelity, we incorporate distribution matching loss (DMD), which constrains generated images to remain within the image manifold learned by pretrained models. We evaluate our method on standard benchmarks and include an extensive ablation study. Without any paired data, our method performs on par with various image editing diffusion models trained on extensive supervised paired data, under the few-step setting. Given the same VLM as the reward model, we also outperform RL-based techniques like Flow-GRPO.

cs.CV

Identifying Prompted Artist Names from Generated Images

A common and controversial use of text-to-image models is to generate pictures by explicitly naming artists, such as "in the style of Greg Rutkowski". We introduce a benchmark for prompted-artist recognition: predicting which artist names were invoked in the prompt from the image alone. The dataset contains 1.95M images covering 110 artists and spans four generalization settings: held-out artists, increasing prompt complexity, multiple-artist prompts, and different text-to-image models. We evaluate feature similarity baselines, contrastive style descriptors, data attribution methods, supervised classifiers, and few-shot prototypical networks. Generalization patterns vary: supervised and few-shot models excel on seen artists and complex prompts, whereas style descriptors transfer better when the artist's style is pronounced; multi-artist prompts remain the most challenging. Our benchmark reveals substantial headroom and provides a public testbed to advance the responsible moderation of text-to-image models. We release the dataset and benchmark to foster further research: https://graceduansu.github.io/IdentifyingPromptedArtists/

cs.CV

Mapping Transient Structures of Cyclo[18]Carbon by Computational X-Ray Spectra

The structure of cyclo[18]carbon (C$_{18}$), whether in its polyynic form with bond length alternation (BLA) or its cumulenic form without BLA, has long fascinated researchers, even prior to its successful synthesis. Recent studies suggest a polyynic ground state and a cumulenic transient state; however, the dynamics remain unclear and lack experimental validation. This study presents a first-principles theoretical investigation of the bond lengths ($R_1$ and $R_2$) dependent two-dimensional potential energy surfaces (PESs) of C$_{18}$, concentrating on the ground state and carbon 1s ionized and excited states. We examine the potential of X-ray spectra for determining bond lengths and monitoring transient structures, finding that both X-ray photoelectron (XPS) and absorption (XAS) spectra are sensitive to these variations. Utilizing a library of ground-state minimum structures optimized with 14 different functionals, we observe that core binding energies predicted with the $\omega$B97XD functional can vary by 0.9 eV (290.3--291.2 eV). Unlike the ground state PES, which predicts minima at alternating bond lengths, the C1s ionized state PES predicts minima with equivalent bond lengths. In the XAS spectra, peaks 1$\pi^*$ and 2$\pi^*$ show a redshift with increasing bond lengths along the line where $R_1 = R_2$. Additionally, increasing $R_2$ (with $R_1$ fixed) results in an initial redshift followed by a blueshift, minimizing at $R_1 = R_2$. Major peaks indicate that both 1$\pi^*$ and 2$\pi^*$ arise from two channels: C1s$\rightarrow\pi^*_{z}$ (out-of-plane) and C1s$\rightarrow\pi^*_{xy}$ (in-plane) transitions at coinciding energies.

physics.chem-ph

Predicting Accurate X-ray Absorption Spectra for CN$^+$, CN, and CN$^-$: Insights from Multiconfigurational and Density Functional Simulations

High-resolution X-ray spectroscopy is an essential tool in X-ray astronomy, enabling detailed studies of celestial objects and their physical and chemical properties. However, comprehensive mapping of high-resolution X-ray spectra for even simple interstellar and circumstellar molecules is still lacking. In this study, we conducted systematic quantum chemical simulations to predict the C1s X-ray absorption spectra of CN$^+$, CN, and CN$^-$. Our findings provide valuable references for both X-ray astronomy and laboratory studies. We assigned the first electronic peak of CN$^+$ and CN to C1s $\rightarrow \sigma^*$ transitions, while the peak for CN$^-$ corresponds to a C1s $\rightarrow \pi^*$ transition. We explained that the two-fold degeneracy ($\pi^*_{xz}$ and $\pi^*_{yz}$) of the C1s$\rightarrow\pi^*$ transitions results in CN$^-$ exhibiting a significantly stronger first absorption compared to the other two systems. We further calculated the vibronic fine structures for these transitions using the quantum wavepacket method based on multiconfigurational-level, anharmonic potential energy curves, revealing distinct energy positions for the 0-0 absorptions at 280.7 eV, 279.6 eV, and 285.8 eV. Each vibronic profile features a prominent 0-0 peak, showing overall similarity but differing intensity ratios of the 0-0 and 0-1 peaks. Notably, introducing a C1s core hole leads to shortened C-N bond lengths and increased vibrational frequencies across all species. These findings enhance our understanding of the electronic structures and X-ray spectra of carbon-nitrogen species, emphasizing the influence of charge state on X-ray absorptions.

physics.chem-ph

Data Attribution for Text-to-Image Models by Unlearning Synthesized Images

The goal of data attribution for text-to-image models is to identify the training images that most influence the generation of a new image. Influence is defined such that, for a given output, if a model is retrained from scratch without the most influential images, the model would fail to reproduce the same output. Unfortunately, directly searching for these influential images is computationally infeasible, since it would require repeatedly retraining models from scratch. In our work, we propose an efficient data attribution method by simulating unlearning the synthesized image. We achieve this by increasing the training loss on the output image, without catastrophic forgetting of other, unrelated concepts. We then identify training images with significant loss deviations after the unlearning process and label these as influential. We evaluate our method with a computationally intensive but "gold-standard" retraining from scratch and demonstrate our method's advantages over previous methods.

cs.CV

Customizing Text-to-Image Models with a Single Image Pair

Art reinterpretation is the practice of creating a variation of a reference work, making a paired artwork that exhibits a distinct artistic style. We ask if such an image pair can be used to customize a generative model to capture the demonstrated stylistic difference. We propose Pair Customization, a new customization method that learns stylistic difference from a single image pair and then applies the acquired style to the generation process. Unlike existing methods that learn to mimic a single concept from a collection of images, our method captures the stylistic difference between paired images. This allows us to apply a stylistic change without overfitting to the specific image content in the examples. To address this new task, we employ a joint optimization method that explicitly separates the style and content into distinct LoRA weight spaces. We optimize these style and content weights to reproduce the style and content images while encouraging their orthogonality. During inference, we modify the diffusion process via a new style guidance based on our learned weights. Both qualitative and quantitative experiments show that our method can effectively learn style while avoiding overfitting to image content, highlighting the potential of modeling such stylistic differences from a single image pair.

cs.CV

Effects of Structural Variations to X-ray Absorption Spectra of g-C$_3$N$_4$: Insights from DFT and TDDFT Simulations

X-ray absorption spectroscopy (XAS) is widely employed for structure characterization of graphitic carbon nitride (g-C$_3$N$_4$) and its composites. Nevertheless, even for pure g-C$_3$N$_4$, discrepancies in energy and profile exist across different experiments, which can be attributed to variations in structures arising from diverse synthesis conditions and calibration procedures. Here, we conducted a theoretical investigation on XAS of three representative g-C$_3$N$_4$ structures (planar, corrugated, and micro-corrugated) optimized with different strategies, to understand the structure-spectroscopy relation. Different methods were compared, including density functional theory (DFT) with the full (FCH) or equivalent (ECH) core-hole approximation, as well as the time-dependent DFT (TDDFT). FCH was responsible for getting accurate absolute absorption energy; while ECH and TDDFT aided in interpreting the spectra, through ECH-state canonical molecular orbitals (ECH-CMOs) and natural transition orbitals (NTOs), respectively. With each method, the spectra at the three structures show evident differences, which can be correlated to different individual experiments or in between. Our calculations explained the structural reason behind the spectral discrepancies among different experiments. Moreover, profiles predicted by these methods also displayed consistency, so their differences can be used as a reliable indicator of their accuracy. Both ECH-CMOs and NTO particle orbitals led to similar graphics, validating their applicability in interpreting the transitions. This work provides a comprehensive analysis of the structure-XAS relation for g-C$_3$N$_4$, provides concrete explanations for the spectral differences reported in various experiments, and offers insight for future structure dynamical and transient X-ray spectral analyses.

cond-mat.mtrl-sci

Global Impact and Balancing Act: Deciphering the Effect of Fluorination on B1s Binding Energies in Fluorinated $h$-BN Nanosheets

X-ray photoelectron spectroscopy (XPS) is an important characterization tool in the pursuit of controllable fluorination of two-dimensional hexagonal boron nitride ($h$-BN). However, there is a lack of clear spectral interpretation and seemingly conflicting measurements exist. To discern the structure-spectroscopy relation, we performed a comprehensive first-principles study on the boron 1s edge XPS of fluorinated $h$-BN (F-BN) nanosheets. By gradually introducing 1--6 fluorine atoms into different boron or nitrogen sites, we created various F-BN structures with doping ratios ranging from 1-6\%. Our calculations reveal that fluorines landed at boron or nitrogen sites exert competitive effects on the B1s binding energies (BEs), leading to red or blue shifts in different measurements. Our calculations affirmed the hypothesis that fluorination affects 1s BEs of all borons in the $\pi$-conjugated system, undermining the transferability from $h$-BN to F-BN. Additionally, we observe that BE generally increases with higher fluorine concentration when both borons and nitrogens are non-exclusively fluorinated. These findings provide critical insights into how fluorination affects boron's 1s BEs, contributing to a better understanding of fluorination functionalization processes in $h$-BN and its potential applications in materials science.

cond-mat.mtrl-sci

Content-Based Search for Deep Generative Models

The growing proliferation of customized and pretrained generative models has made it infeasible for a user to be fully cognizant of every model in existence. To address this need, we introduce the task of content-based model search: given a query and a large set of generative models, finding the models that best match the query. As each generative model produces a distribution of images, we formulate the search task as an optimization problem to select the model with the highest probability of generating similar content as the query. We introduce a formulation to approximate this probability given the query from different modalities, e.g., image, sketch, and text. Furthermore, we propose a contrastive learning framework for model retrieval, which learns to adapt features for various query modalities. We demonstrate that our method outperforms several baselines on Generative Model Zoo, a new benchmark we create for the model retrieval task.

cs.CV

Ablating Concepts in Text-to-Image Diffusion Models

Large-scale text-to-image diffusion models can generate high-fidelity images with powerful compositional ability. However, these models are typically trained on an enormous amount of Internet data, often containing copyrighted material, licensed images, and personal photos. Furthermore, they have been found to replicate the style of various living artists or memorize exact training samples. How can we remove such copyrighted concepts or images without retraining the model from scratch? To achieve this goal, we propose an efficient method of ablating concepts in the pretrained model, i.e., preventing the generation of a target concept. Our algorithm learns to match the image distribution for a target style, instance, or text prompt we wish to ablate to the distribution corresponding to an anchor concept. This prevents the model from generating target concepts given its text condition. Extensive experiments show that our method can successfully prevent the generation of the ablated concept while preserving closely related concepts in the model.

cs.CV

Evaluating Data Attribution for Text-to-Image Models

While large text-to-image models are able to synthesize "novel" images, these images are necessarily a reflection of the training data. The problem of data attribution in such models -- which of the images in the training set are most responsible for the appearance of a given generated image -- is a difficult yet important one. As an initial step toward this problem, we evaluate attribution through "customization" methods, which tune an existing large-scale model toward a given exemplar object or style. Our key insight is that this allows us to efficiently create synthetic images that are computationally influenced by the exemplar by construction. With our new dataset of such exemplar-influenced images, we are able to evaluate various data attribution algorithms and different possible feature spaces. Furthermore, by training on our dataset, we can tune standard models, such as DINO, CLIP, and ViT, toward the attribution problem. Even though the procedure is tuned towards small exemplar sets, we show generalization to larger sets. Finally, by taking into account the inherent uncertainty of the problem, we can assign soft attribution scores over a set of training images.

cs.CV

Rewriting Geometric Rules of a GAN

Deep generative models make visual content creation more accessible to novice users by automating the synthesis of diverse, realistic content based on a collected dataset. However, the current machine learning approaches miss a key element of the creative process -- the ability to synthesize things that go far beyond the data distribution and everyday experience. To begin to address this issue, we enable a user to "warp" a given model by editing just a handful of original model outputs with desired geometric changes. Our method applies a low-rank update to a single model layer to reconstruct edited examples. Furthermore, to combat overfitting, we propose a latent space augmentation method based on style-mixing. Our method allows a user to create a model that synthesizes endless objects with defined geometric changes, enabling the creation of a new generative model without the burden of curating a large-scale dataset. We also demonstrate that edited models can be composed to achieve aggregated effects, and we present an interactive interface to enable users to create new models through composition. Empirical measurements on multiple test cases suggest the advantage of our method against recent GAN fine-tuning methods. Finally, we showcase several applications using the edited models, including latent space interpolation and image editing.

cs.CV

Sketch Your Own GAN

Can a user create a deep generative model by sketching a single example? Traditionally, creating a GAN model has required the collection of a large-scale dataset of exemplars and specialized knowledge in deep learning. In contrast, sketching is possibly the most universally accessible way to convey a visual concept. In this work, we present a method, GAN Sketching, for rewriting GANs with one or more sketches, to make GANs training easier for novice users. In particular, we change the weights of an original GAN model according to user sketches. We encourage the model's output to match the user sketches through a cross-domain adversarial loss. Furthermore, we explore different regularization methods to preserve the original model's diversity and image quality. Experiments have shown that our method can mold GANs to match shapes and poses specified by sketches while maintaining realism and diversity. Finally, we demonstrate a few applications of the resulting GAN, including latent space interpolation and image editing.

cs.CV

CNN-generated images are surprisingly easy to spot... for now

In this work we ask whether it is possible to create a "universal" detector for telling apart real images from these generated by a CNN, regardless of architecture or dataset used. To test this, we collect a dataset consisting of fake images generated by 11 different CNN-based image generator models, chosen to span the space of commonly used architectures today (ProGAN, StyleGAN, BigGAN, CycleGAN, StarGAN, GauGAN, DeepFakes, cascaded refinement networks, implicit maximum likelihood estimation, second-order attention super-resolution, seeing-in-the-dark). We demonstrate that, with careful pre- and post-processing and data augmentation, a standard image classifier trained on only one specific CNN generator (ProGAN) is able to generalize surprisingly well to unseen architectures, datasets, and training methods (including the just released StyleGAN2). Our findings suggest the intriguing possibility that today's CNN-generated images share some common systematic flaws, preventing them from achieving realistic image synthesis. Code and pre-trained networks are available at https://peterwang512.github.io/CNNDetection/ .

cs.CV

Detecting Photoshopped Faces by Scripting Photoshop

Most malicious photo manipulations are created using standard image editing tools, such as Adobe Photoshop. We present a method for detecting one very popular Photoshop manipulation -- image warping applied to human faces -- using a model trained entirely using fake images that were automatically generated by scripting Photoshop itself. We show that our model outperforms humans at the task of recognizing manipulated images, can predict the specific location of edits, and in some cases can be used to "undo" a manipulation to reconstruct the original, unedited image. We demonstrate that the system can be successfully applied to real, artist-created image manipulations.

cs.CV