SearcharxivSearch

arXiv subjects

Yue Gong

Publications and source records attributed to Yue Gong.

15 recordsLinked to original sources

RPiAE: A Representation-Pivoted Autoencoder Enhancing Both Image Generation and Editing

Diffusion models have become the dominant paradigm for image generation and editing, with latent diffusion models shifting denoising to a compact latent space for efficiency and scalability. Recent attempts to leverage pretrained visual representation models as tokenizer priors either align diffusion features to representation features or directly reuse representation encoders as frozen tokenizers. Although such approaches can improve generation metrics, they often suffer from limited reconstruction fidelity due to frozen encoders, which in turn degrades editing quality, as well as overly high-dimensional latents that make diffusion modeling difficult. To address these limitations, We propose Representation-Pivoted AutoEncoder, a representation-based tokenizer that improves both generation and editing. We introduce Representation-Pivot Regularization, a training strategy that enables a representation-initialized encoder to be fine-tuned for reconstruction while preserving the semantic structure of the pretrained representation space, followed by a variational bridge which compress latent space into a compact one for better diffusion modeling. We adopt an objective-decoupled stage-wise training strategy that sequentially optimizes generative tractability and reconstruction-fidelity objectives. Together, these components yield a tokenizer that preserves strong semantics, reconstructs faithfully, and produces latents with reduced diffusion modeling complexity. Experiments demonstrate that RPiAE outperforms other visual tokenizers on text-to-image generation and image editing, while delivering the best reconstruction fidelity among representation-based tokenizers.

cs.CV

RefTon: Reference person shot assist virtual Try-on

We introduce RefTon, a flux-based person-to-person virtual try-on framework that enhances garment realism through unpaired visual references. Unlike conventional approaches that rely on complex auxiliary inputs such as body parsing and warped mask or require finely designed extract branches to process various input conditions, RefTon streamlines the process by directly generating try-on results from a source image and a target garment, without the need for structural guidance or auxiliary components to handle diverse inputs. Moreover, inspired by human clothing selection behavior, RefTon leverages additional reference images (the target garment worn on different individuals) to provide powerful guidance for refining texture alignment and maintaining the garment details. To enable this capability, we built a dataset containing unpaired reference images for training. Extensive experiments on public benchmarks demonstrate that RefTon achieves competitive or superior performance compared to state-of-the-art methods, while maintaining a simple and efficient person-to-person design.

cs.CV

Asymmetric stress engineering of dense dislocations in brittle superconductors for strong vortex pinning

Large lossless currents in high-temperature superconductors (HTS) critically rely on dense defects with suitable size and dimensionality to pin vortices, with dislocations being particularly effective due to their one-dimensional geometry to interact extensively with vortex lines. However, in non-metallic compounds such as HTS with rigid lattices, conventional deformation methods typically lead to catastrophic fracture rather than dislocation-mediated plasticity, making it a persistent challenge to introduce dislocations at high density. Here, we propose an asymmetric stress field strategy using extrusion to directly nucleate a high-density of dislocations in HTS by activating shear-driven lattice slip and twisting under superimposed hydrostatic compression. As demonstrated in iron-based superconductors (IBS), atomic displacements of nearly one angstrom trigger the formation of tilted dislocation lines with a density approaching that of metals. With further structural refinement, these dislocations serve as strong pinning centers that lead to a fivefold enhancement in the current-carrying capacity of IBS at 33 T, along with low anisotropy and a large irreversibility field. This work not only establishes a scalable route to engineer pinning landscapes in HTS, but also offers a generalizable framework for manipulating dislocation structures in rigid crystalline systems.

cond-mat.supr-con

CTA-Flux: Integrating Chinese Cultural Semantics into High-Quality English Text-to-Image Communities

We proposed the Chinese Text Adapter-Flux (CTA-Flux). An adaptation method fits the Chinese text inputs to Flux, a powerful text-to-image (TTI) generative model initially trained on the English corpus. Despite the notable image generation ability conditioned on English text inputs, Flux performs poorly when processing non-English prompts, particularly due to linguistic and cultural biases inherent in predominantly English-centric training datasets. Existing approaches, such as translating non-English prompts into English or finetuning models for bilingual mappings, inadequately address culturally specific semantics, compromising image authenticity and quality. To address this issue, we introduce a novel method to bridge Chinese semantic understanding with compatibility in English-centric TTI model communities. Existing approaches relying on ControlNet-like architectures typically require a massive parameter scale and lack direct control over Chinese semantics. In comparison, CTA-flux leverages MultiModal Diffusion Transformer (MMDiT) to control the Flux backbone directly, significantly reducing the number of parameters while enhancing the model's understanding of Chinese semantics. This integration significantly improves the generation quality and cultural authenticity without extensive retraining of the entire model, thus maintaining compatibility with existing text-to-image plugins such as LoRA, IP-Adapter, and ControlNet. Empirical evaluations demonstrate that CTA-flux supports Chinese and English prompts and achieves superior image generation quality, visual realism, and faithful depiction of Chinese semantics.

cs.CV

NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

Diffusion Transformers (DiTs) have demonstrated exceptional capabilities in text-to-image synthesis. However, in the domain of controllable text-to-image generation using DiTs, most existing methods still rely on the ControlNet paradigm originally designed for UNet-based diffusion models. This paradigm introduces significant parameter overhead and increased computational costs. To address these challenges, we propose the Nano Control Diffusion Transformer (NanoControl), which employs Flux as the backbone network. Our model achieves state-of-the-art controllable text-to-image generation performance while incurring only a 0.024\% increase in parameter count and a 0.029\% increase in GFLOPs, thus enabling highly efficient controllable generation. Specifically, rather than duplicating the DiT backbone for control, we design a LoRA-style (low-rank adaptation) control module that directly learns control signals from raw conditioning inputs. Furthermore, we introduce a KV-Context Augmentation mechanism that integrates condition-specific key-value information into the backbone in a simple yet highly effective manner, facilitating deep fusion of conditional features. Extensive benchmark experiments demonstrate that NanoControl significantly reduces computational overhead compared to conventional control approaches, while maintaining superior generation quality and achieving improved controllability.

cs.CV

FLUX-Makeup: High-Fidelity, Identity-Consistent, and Robust Makeup Transfer via Diffusion Transformer

Makeup transfer aims to apply the makeup style from a reference face to a target face and has been increasingly adopted in practical applications. Existing GAN-based approaches typically rely on carefully designed loss functions to balance transfer quality and facial identity consistency, while diffusion-based methods often depend on additional face-control modules or algorithms to preserve identity. However, these auxiliary components tend to introduce extra errors, leading to suboptimal transfer results. To overcome these limitations, we propose FLUX-Makeup, a high-fidelity, identity-consistent, and robust makeup transfer framework that eliminates the need for any auxiliary face-control components. Instead, our method directly leverages source-reference image pairs to achieve superior transfer performance. Specifically, we build our framework upon FLUX-Kontext, using the source image as its native conditional input. Furthermore, we introduce RefLoRAInjector, a lightweight makeup feature injector that decouples the reference pathway from the backbone, enabling efficient and comprehensive extraction of makeup-related information. In parallel, we design a robust and scalable data generation pipeline to provide more accurate supervision during training. The paired makeup datasets produced by this pipeline significantly surpass the quality of all existing datasets. Extensive experiments demonstrate that FLUX-Makeup achieves state-of-the-art performance, exhibiting strong robustness across diverse scenarios.

cs.CV

SQLens: An End-to-End Framework for Error Detection and Correction in Text-to-SQL

Text-to-SQL systems translate natural language (NL) questions into SQL queries, enabling non-technical users to interact with structured data. While large language models (LLMs) have shown promising results on the text-to-SQL task, they often produce semantically incorrect yet syntactically valid queries, with limited insight into their reliability. We propose SQLens, an end-to-end framework for fine-grained detection and correction of semantic errors in LLM-generated SQL. SQLens integrates error signals from both the underlying database and the LLM to identify potential semantic errors within SQL clauses. It further leverages these signals to guide query correction. Empirical results on two public benchmarks show that SQLens outperforms the best LLM-based self-evaluation method by 25.78% in F1 for error detection, and improves execution accuracy of out-of-the-box text-to-SQL systems by up to 20%.

cs.CL

Exploiting LLMs for Automatic Hypothesis Assessment via a Logit-Based Calibrated Prior

As hypothesis generation becomes increasingly automated, a new bottleneck has emerged: hypothesis assessment. Modern systems can surface thousands of statistical relationships-correlations, trends, causal links-but offer little guidance on which ones are novel, non-trivial, or worthy of expert attention. In this work, we study the complementary problem to hypothesis generation: automatic hypothesis assessment. Specifically, we ask: given a large set of statistical relationships, can we automatically assess which ones are novel and worth further exploration? We focus on correlations as they are a common entry point in exploratory data analysis that often serve as the basis for forming deeper scientific or causal hypotheses. To support automatic assessment, we propose to leverage the vast knowledge encoded in LLMs' weights to derive a prior distribution over the correlation value of a variable pair. If an LLM's prior expects the correlation value observed, then such correlation is not surprising, and vice versa. We propose the Logit-based Calibrated Prior, an LLM-elicited correlation prior that transforms the model's raw output logits into a calibrated, continuous predictive distribution over correlation values. We evaluate the prior on a benchmark of 2,096 real-world variable pairs and it achieves a sign accuracy of 78.8%, a mean absolute error of 0.26, and 95% credible interval coverage of 89.2% in predicting Pearson correlation coefficient. It also outperforms a fine-tuned RoBERTa classifier in binary correlation prediction and achieves higher precision@K in hypothesis ranking. We further show that the prior generalizes to correlations not seen during LLM pretraining, reflecting context-sensitive reasoning rather than memorization.

cs.LG

Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System

Finding relevant tables among databases, lakes, and repositories is the first step in extracting value from data. Such a task remains difficult because assessing whether a table is relevant to a problem does not always depend only on its content but also on the context, which is usually tribal knowledge known to the individual or team. While tools like data catalogs and academic data discovery systems target this problem, they rely on keyword search or more complex interfaces, limiting non-technical users' ability to find relevant data. The advent of large language models (LLMs) offers a unique opportunity for users to ask questions directly in natural language, making dataset discovery more intuitive, accessible, and efficient. In this paper, we introduce Pneuma, a retrieval-augmented generation (RAG) system designed to efficiently and effectively discover tabular data. Pneuma leverages large language models (LLMs) for both table representation and table retrieval. For table representation, Pneuma preserves schema and row-level information to ensure comprehensive data understanding. For table retrieval, Pneuma augments LLMs with traditional information retrieval techniques, such as full-text and vector search, harnessing the strengths of both to improve retrieval performance. To evaluate Pneuma, we generate comprehensive benchmarks that simulate table discovery workload on six real-world datasets including enterprise data, scientific databases, warehousing data, and open data. Our results demonstrate that Pneuma outperforms widely used table search systems (such as full-text search and state-of-the-art RAG systems) in accuracy and resource efficiency.

cs.DB

A survey on deep learning approaches for data integration in autonomous driving system

The perception module of self-driving vehicles relies on a multi-sensor system to understand its environment. Recent advancements in deep learning have led to the rapid development of approaches that integrate multi-sensory measurements to enhance perception capabilities. This paper surveys the latest deep learning integration techniques applied to the perception module in autonomous driving systems, categorizing integration approaches based on "what, how, and when to integrate". A new taxonomy of integration is proposed, based on three dimensions: multi-view, multi-modality, and multi-frame. The integration operations and their pros and cons are summarized, providing new insights into the properties of an "ideal" data integration approach that can alleviate the limitations of existing methods. After reviewing hundreds of relevant papers, this survey concludes with a discussion of the key features of an optimal data integration approach.

cs.RO

METAM: Goal-Oriented Data Discovery

Data is a central component of machine learning and causal inference tasks. The availability of large amounts of data from sources such as open data repositories, data lakes and data marketplaces creates an opportunity to augment data and boost those tasks' performance. However, augmentation techniques rely on a user manually discovering and shortlisting useful candidate augmentations. Existing solutions do not leverage the synergy between discovery and augmentation, thus under exploiting data. In this paper, we introduce METAM, a novel goal-oriented framework that queries the downstream task with a candidate dataset, forming a feedback loop that automatically steers the discovery and augmentation process. To select candidates efficiently, METAM leverages properties of the: i) data, ii) utility function, and iii) solution set size. We show METAM's theoretical guarantees and demonstrate those empirically on a broad set of tasks. All in all, we demonstrate the promise of goal-oriented data discovery to modern data science applications.

cs.DB

Ver: View Discovery in the Wild

We present Ver, a data discovery system that identifies project-join views over large repositories of tables that do not contain join path information, and even when input queries are inaccurate. Ver implements a reference architecture to solve both the technical (scale and search) and human (semantic ambiguity, navigating a large number of results) problems of view discovery. We demonstrate users find the view they want when using Ver with a user study and we demonstrate its performance with large-scale end-to-end experiments on real-world datasets containing tens of millions of join paths.

cs.DB

Two-dimensional spinodal interface in one-step grown graphene-molybdenum carbide heterostructures

Heterostructures made by stacking different materials on top of each other are expected to exhibit unusual properties and new phenomena. Interface of the heterostructures plays a vital role in determining their properties. Here, we report the observation of a two-dimensional (2D) spinodal interface in graphene-molybdenum carbide (α-Mo2C) heterostructures, which arises from spinodal decomposition occurring at the heterointerface, by using scanning tunneling microscopy. Our experiment demonstrates that the 2D spinodal interface modulates graphene into whispering gallery resonant networks filled with quasi-bound states of massless Dirac fermions. Moreover, below the superconducting transition temperature of the underlying α-Mo2C, the 2D spinodal interface behaves as disorders, resulting in the breakdown of the proximity-induced superconductivity in graphene. Our result sheds new light on tuning properties of heterostructures based on interface engineering.

cond-mat.mtrl-sci

Metallic vanadium disulfide nanosheets as a platform material for multifunctional electrode applications

Nano-thick metallic transition metal dichalcogenides such as VS$_{2}$ are essential building blocks for constructing next-generation electronic and energy-storage applications, as well as for exploring unique physical issues associated with the dimensionality effect. However, such 2D layered materials have yet to be achieved through either mechanical exfoliation or bottom-up synthesis. Herein, we report a facile chemical vapor deposition route for direct production of crystalline VS$_{2}$ nanosheets with sub-10 nm thicknesses and domain sizes of tens of micrometers. The obtained nanosheets feature spontaneous superlattice periodicities and excellent electrical conductivities (~3$\times$10$^{3}$ S cm$^{-1}$), which has enabled a variety of applications such as contact electrodes for monolayer MoS$_{2}$ with contact resistances of ~1/4 to that of Ni/Au metals, and as supercapacitor electrodes in aqueous electrolytes showing specific capacitances as high as 8.6$\times$10$^{2}$ F g$^{-1}$. This work provides fresh insights into the delicate structure-property relationship and the broad application prospects of such metallic 2D materials.

cond-mat.mtrl-sci

One-step synthesis of van der Waals heterostructures of graphene and 2D superconducting a-Mo2C

Assembling different two-dimensional (2D) crystals, covering a very broad range of properties, into van der Waals (vdW) heterostructures enables the unprecedented possibilities for combining the best of different ingredients in one objective material. So far, metallic, semiconducting, and insulating 2D crystals have been used successfully in making functional vdW heterostructures with properties by design. Here, we expand 2D superconducting crystals as a building block of the vdW hererostructures. A one-step growth of large-scale high-quality vdW heterostructures of graphene and 2D superconducting a-Mo2C by using chemical vapor deposition (CVD) method is reported. The superconductivity and its 2D nature of the heterostructures are characterized by our scanning tunneling microscopy (STM) measurements. This adds the 2D superconductivity, the most attractive property of condensed matter physics, to the vdW heterostructures.

cond-mat.mtrl-sci