SearcharxivSearch

arXiv subjects

Chengyuan Xu

Publications and source records attributed to Chengyuan Xu.

7 recordsLinked to original sources

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes. Existing approaches typically train individual adapter for each attribute separately, overlooking the entanglement among multiple attributes. As a result, interference occurs among different attributes, preventing precise control of multiple attributes together. To address this challenge, we aim to disentangle multiple attributes in slider-based generation to enbale more reliable and independent attribute manipulation. Our approach, CompSlider, can generate a conditional prior for the T2I foundation model to control multiple attributes simultaneously. Furthermore, we introduce novel disentanglement and structure losses to compose multiple attribute changes while maintaining structural consistency within the image. Since CompSlider operates in the latent space of the conditional prior and does not require retraining the foundation model, it reduces the computational burden for both training and inference. We evaluate our approach on a variety of image attributes and highlight its generality by extending to video generation.

cs.CV

Reduced model for deep bed filtration of binary highly heterogeneous colloids

Hyper-exponential retention profiles (HERPs) are often observed during laboratory tests on colloidal and nano-suspension transport in porous media. The aim of this work is an extension of the traditional model for suspended particle transport (deep bed filtration) to match HERPs. Interpreting particle capture in the framework of active mass law for the "chemical reaction" particle-rock, we substitute suspension concentration in the traditional expression for capture rate by non-linear "activity" function, called the suspension function. The non-linear suspension function is derived from Margules' activity model for excess Gibbs free energy. This suspension function is shown to match that derived from multicomponent colloidal transport with good agreement, despite having one less unknown parameter. Empirical formulae are presented to convert between the two unknown parameters of the multicomponent model and the Margules exponent. Comparison between the new model and laboratory coreflooding data shows good agreement, even in cases with distinctly hyper-exponential retention profiles, where the classical model fails to reproduce the laboratory data.

physics.geo-ph

Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI

Seamless integration of virtual and physical worlds in augmented reality benefits from the system semantically "understanding" the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel capabilities that leverage the semantics in the 3D environment for various object-level interactions. Meanwhile, the computer vision community has made leaps in neural vision-language understanding to enhance environment perception for autonomous tasks. In this work, we introduce a multimodal 3D object representation that unifies both semantic and linguistic knowledge with the geometric representation, enabling user-guided machine learning involving physical objects. We first present a fast multimodal 3D reconstruction pipeline that brings linguistic understanding to AR by fusing CLIP vision-language features into the environment and object models. We then propose "in-situ" machine learning, which, in conjunction with the multimodal representation, enables new tools and interfaces for users to interact with physical spaces and objects in a spatially and linguistically meaningful manner. We demonstrate the usefulness of the proposed system through two real-world AR applications on Magic Leap 2: a) spatial search in physical environments with natural language and b) an intelligent inventory system that tracks object changes over time. We also make our full implementation and demo data available at (https://github.com/cy-xu/spatially_aware_AI) to encourage further exploration and research in spatially aware AI.

cs.HC

Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task

When designing an AI-assisted decision-making system, there is often a tradeoff between precision and recall in the AI's recommendations. We argue that careful exploitation of this tradeoff can harness the complementary strengths in the human-AI collaboration to significantly improve team performance. We investigate a real-world video anonymization task for which recall is paramount and more costly to improve. We analyze the performance of 78 professional annotators working with a) no AI assistance, b) a high-precision "restrained" AI, and c) a high-recall "zealous" AI in over 3,466 person-hours of annotation work. In comparison, the zealous AI helps human teammates achieve significantly shorter task completion time and higher recall. In a follow-up study, we remove AI assistance for everyone and find negative training effects on annotators trained with the restrained AI. These findings and our analysis point to important implications for the design of AI assistance in recall-demanding scenarios.

cs.HC

Cosmic-CoNN: A Cosmic Ray Detection Deep-Learning Framework, Dataset, and Toolkit

Rejecting cosmic rays (CRs) is essential for the scientific interpretation of CCD-captured data, but detecting CRs in single-exposure images has remained challenging. Conventional CR detectors require experimental parameter tuning for different instruments, and recent deep learning methods only produce instrument-specific models that suffer from performance loss on telescopes not included in the training data. We present Cosmic-CoNN, a generic CR detector deployed for 24 telescopes at the Las Cumbres Observatory, which is made possible by the three contributions in this work: 1) We build a large and diverse ground-based CR dataset leveraging thousands of images from a global telescope network. 2) We propose a novel loss function and a neural network optimized for telescope imaging data to train generic CR detection models. At 95% recall, our model achieves a precision of 93.70% on Las Cumbres imaging data and maintains a consistent performance on new ground-based instruments never used for training. Specifically, the Cosmic-CoNN model trained on the Las Cumbres CR dataset maintains high precisions of 92.03% and 96.69% on Gemini GMOS-N/S 1x1 and 2x2 binning images, respectively. 3) We build a suite of tools including an interactive CR mask visualization and editing interface, console commands, and Python APIs to make automatic, robust CR detection widely accessible by the community of astronomers. Our dataset, open-source codebase, and trained models are available at https://github.com/cy-xu/cosmic-conn.

astro-ph.IM

Interactive Segmentation and Visualization for Tiny Objects in Multi-megapixel Images

We introduce an interactive image segmentation and visualization framework for identifying, inspecting, and editing tiny objects (just a few pixels wide) in large multi-megapixel high-dynamic-range (HDR) images. Detecting cosmic rays (CRs) in astronomical observations is a cumbersome workflow that requires multiple tools, so we developed an interactive toolkit that unifies model inference, HDR image visualization, segmentation mask inspection and editing into a single graphical user interface. The feature set, initially designed for astronomical data, makes this work a useful research-supporting tool for human-in-the-loop tiny-object segmentation in scientific areas like biomedicine, materials science, remote sensing, etc., as well as computer vision. Our interface features mouse-controlled, synchronized, dual-window visualization of the image and the segmentation mask, a critical feature for locating tiny objects in multi-megapixel images. The browser-based tool can be readily hosted on the web to provide multi-user access and GPU acceleration for any device. The toolkit can also be used as a high-precision annotation tool, or adapted as the frontend for an interactive machine learning framework. Our open-source dataset, CR detection model, and visualization toolkit are available at https://github.com/cy-xu/cosmic-conn.

cs.CV

The electron-capture origin of supernova 2018zd

In the transitional mass range ($\sim$ 8-10 solar masses) between white dwarf formation and iron core-collapse supernovae, stars are expected to produce an electron-capture supernova. Theoretically, these progenitors are thought to be super-asymptotic giant branch stars with a degenerate O+Ne+Mg core, and electron capture onto Ne and Mg nuclei should initiate core collapse. However, no supernovae have unequivocally been identified from an electron-capture origin, partly because of uncertainty in theoretical predictions. Here we present six indicators of electron-capture supernovae and show that supernova 2018zd is the only known supernova having strong evidence for or consistent with all six: progenitor identification, circumstellar material, chemical composition, explosion energy, light curve, and nucleosynthesis. For supernova 2018zd, we infer a super-asymptotic giant branch progenitor based on the faint candidate in the pre-explosion images and the chemically-enriched circumstellar material revealed by the early ultraviolet colours and flash spectroscopy. The light-curve morphology and nebular emission lines can be explained with the low explosion energy and neutron-rich nucleosynthesis produced in an electron-capture supernova. This identification provides insights into the complex stellar evolution, supernova physics, cosmic nucleosynthesis, and remnant populations in the transitional mass range.

astro-ph.HE