Searcharxiv⌕ Search

arXiv subjects

Chong Sun

Publications and source records attributed to Chong Sun.

At least 37 records · Page 2Linked to original sources

Towards Efficient Multi-Scale Deformable Attention on NPU

Multi-scale deformable attention (MSDA) is a flexible and powerful feature extraction mechanism for visual tasks, but its random-access grid sampling strategy poses significant optimization challenges, especially on domain-specific accelerators such as NPUs. In this work, we present a co-design approach that systematically rethinks memory access and computation strategies for MSDA on the Ascend NPU architecture. With this co-design approach, our implementation supports both efficient forward and backward computation, is fully adapted for training workloads, and incorporates a suite of hardware-aware optimizations. Extensive experiments show that our solution achieves up to $5.9\times$ (forward), $8.9\times$ (backward), and $7.3\times$ (end-to-end training) speedup over the grid sample-based baseline, and $1.9\times$, $2.4\times$, and $2.0\times$ acceleration over the latest vendor library, respectively.

cs.PF↗

WeGen: A Unified Model for Interactive Multimodal Generation as We Chat

Existing multimodal generative models fall short as qualified design copilots, as they often struggle to generate imaginative outputs once instructions are less detailed or lack the ability to maintain consistency with the provided references. In this work, we introduce WeGen, a model that unifies multimodal generation and understanding, and promotes their interplay in iterative generation. It can generate diverse results with high creativity for less detailed instructions. And it can progressively refine prior generation results or integrating specific contents from references following the instructions in its chat with users. During this process, it is capable of preserving consistency in the parts that the user is already satisfied with. To this end, we curate a large-scale dataset, extracted from Internet videos, containing rich object dynamics and auto-labeled dynamics descriptions by advanced foundation models to date. These two information are interleaved into a single sequence to enable WeGen to learn consistency-aware generation where the specified dynamics are generated while the consistency of unspecified content is preserved aligned with instructions. Besides, we introduce a prompt self-rewriting mechanism to enhance generation diversity. Extensive experiments demonstrate the effectiveness of unifying multimodal understanding and generation in WeGen and show it achieves state-of-the-art performance across various visual generation benchmarks. These also demonstrate the potential of WeGen as a user-friendly design copilot as desired. The code and models will be available at https://github.com/hzphzp/WeGen.

cs.CV↗

Get In Video: Add Anything You Want to the Video

Video editing increasingly demands the ability to incorporate specific real-world instances into existing footage, yet current approaches fundamentally fail to capture the unique visual characteristics of particular subjects and ensure natural instance/scene interactions. We formalize this overlooked yet critical editing paradigm as "Get-In-Video Editing", where users provide reference images to precisely specify visual elements they wish to incorporate into videos. Addressing this task's dual challenges, severe training data scarcity and technical challenges in maintaining spatiotemporal coherence, we introduce three key contributions. First, we develop GetIn-1M dataset created through our automated Recognize-Track-Erase pipeline, which sequentially performs video captioning, salient instance identification, object detection, temporal tracking, and instance removal to generate high-quality video editing pairs with comprehensive annotations (reference image, tracking mask, instance prompt). Second, we present GetInVideo, a novel end-to-end framework that leverages a diffusion transformer architecture with 3D full attention to process reference images, condition videos, and masks simultaneously, maintaining temporal coherence, preserving visual identity, and ensuring natural scene interactions when integrating reference objects into videos. Finally, we establish GetInBench, the first comprehensive benchmark for Get-In-Video Editing scenario, demonstrating our approach's superior performance through extensive evaluations. Our work enables accessible, high-quality incorporation of specific real-world subjects into videos, significantly advancing personalized video editing capabilities.

cs.CV↗

Taking A Closer Look at Interacting Objects: Interaction-Aware Open Vocabulary Scene Graph Generation

Today's open vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Most existing methods adopt a two-stage pipeline: weakly supervised pre-training with image captions and supervised fine-tuning (SFT) on fully annotated scene graphs. Nonetheless, they omit explicit modeling of interacting objects and treat all objects equally, resulting in mismatched relation pairs. To this end, we propose an interaction-aware OVSGG framework INOVA. During pre-training, INOVA employs an interaction-aware target generation strategy to distinguish interacting objects from non-interacting ones. In SFT, INOVA devises an interaction-guided query selection tactic to prioritize interacting objects during bipartite graph matching. Besides, INOVA is equipped with an interaction-consistent knowledge distillation to enhance the robustness by pushing interacting object pairs away from the background. Extensive experiments on two benchmarks (VG and GQA) show that INOVA achieves state-of-the-art performance, demonstrating the potential of interaction-aware mechanisms for real-world applications.

cs.CV↗

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Video-to-music generation presents significant potential in video production, requiring the generated music to be both semantically and rhythmically aligned with the video. Achieving this alignment demands advanced music generation capabilities, sophisticated video understanding, and an efficient mechanism to learn the correspondence between the two modalities. In this paper, we propose VidMusician, a parameter-efficient video-to-music generation framework built upon text-to-music models. VidMusician leverages hierarchical visual features to ensure semantic and rhythmic alignment between video and music. Specifically, our approach utilizes global visual features as semantic conditions and local visual features as rhythmic cues. These features are integrated into the generative backbone via cross-attention and in-attention mechanisms, respectively. Through a two-stage training process, we incrementally incorporate semantic and rhythmic features, utilizing zero initialization and identity initialization to maintain the inherent music-generative capabilities of the backbone. Additionally, we construct a diverse video-music dataset, DVMSet, encompassing various scenarios, such as promo videos, commercials, and compilations. Experiments demonstrate that VidMusician outperforms state-of-the-art methods across multiple evaluation metrics and exhibits robust performance on AI-generated videos. Samples are available at \url{https://youtu.be/EPOSXwtl1jw}.

cs.SD↗

Waveflow: boundary-conditioned normalizing flows applied to fermionic wavefunctions

An efficient and expressive wavefunction ansatz is key to scalable solutions for complex many-body electronic structures. While Slater determinants are predominantly used for constructing antisymmetric electronic wavefunction ansätze, this construction can result in limited expressiveness when the targeted wavefunction is highly complex. In this work, we introduce Waveflow, an innovative framework for learning many-body fermionic wavefunctions using boundary-conditioned normalizing flows. Instead of relying on Slater determinants, Waveflow imposes antisymmetry by defining the fundamental domain of the wavefunction and applying necessary boundary conditions. A key challenge in using normalizing flows for this purpose is addressing the topological mismatch between the prior and target distributions. We propose using O-spline priors and I-spline bijections to handle this mismatch, which allows for flexibility in the node number of the distribution while automatically maintaining its square-normalization property. We apply Waveflow to a one-dimensional many-electron system, where we variationally minimize the system's energy using variational quantum Monte Carlo (VQMC). Our experiments demonstrate that Waveflow can effectively resolve topological mismatches and faithfully learn the ground-state wavefunction.

cs.LG↗

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks, such as MathVista and MathVerse, focus more on the result-oriented performance but neglect the underlying principles in knowledge acquisition and generalization. Inspired by human-like mathematical reasoning, we introduce WE-MATH, the first benchmark specifically designed to explore the problem-solving principles beyond end-to-end performance. We meticulously collect and categorize 6.5K visual math problems, spanning 67 hierarchical knowledge concepts and five layers of knowledge granularity. We decompose composite problems into sub-problems according to the required knowledge concepts and introduce a novel four-dimensional metric, namely Insufficient Knowledge (IK), Inadequate Generalization (IG), Complete Mastery (CM), and Rote Memorization (RM), to hierarchically assess inherent issues in LMMs' reasoning process. With WE-MATH, we conduct a thorough evaluation of existing LMMs in visual mathematical reasoning and reveal a negative correlation between solving steps and problem-specific performance. We confirm the IK issue of LMMs can be effectively improved via knowledge augmentation strategies. More notably, the primary challenge of GPT-4o has significantly transitioned from IK to IG, establishing it as the first LMM advancing towards the knowledge generalization stage. In contrast, other LMMs exhibit a marked inclination towards Rote Memorization - they correctly solve composite problems involving multiple knowledge concepts yet fail to answer sub-problems. We anticipate that WE-MATH will open new pathways for advancements in visual mathematical reasoning for LMMs. The WE-MATH data and evaluation code are available at https://github.com/We-Math/We-Math.

cs.AI↗

Electron localization in disordered quantum systems at finite temperatures

We study electron localization in disordered quantum systems, focusing on both individual eigenstates and thermal states. We employ complex polarization as a numerical indicator to characterize the system's localization length. Furthermore, we assess the efficacy of mean-field approximation in providing a quantitative analysis of such systems. Through this study, we seek to provide insight into the following aspects: the behavior of electron localization as a function of interaction, disorder, and temperature, whether thermal states and highly excited states exhibit similar properties in many-body localized systems, and the reliability of the mean-field approximation in weak-interaction scenarios.

cond-mat.dis-nn↗

Selected non-orthogonal configuration interaction with compressed single and double excitations

Addressing both dynamic and static correlation accurately is a primary goal in electronic structure theory. Non-orthogonal configuration interaction (NOCI) is a versatile tool for treating static correlation, offering chemical insights by combining diverse reference states. Nevertheless, achieving quantitative accuracy requires the inclusion of missing dynamic correlation. This work introduces a framework for compressing orthogonal single and double excitations into an NOCI of a much smaller dimension. This compression is repeated with each Slater determinant in a reference NOCI, resulting in another NOCI that includes all its single and double excitations (NOCISD), effectively recovering the missing dynamic correlations from the reference. This compressed NOCISD is further refined through a selection process using metric and energy tests (SNOCISD). We validate the effectiveness of SNOCISD through its application to the dissociation of the nitrogen molecule and the hole-doped two-dimensional Hubbard model at various interaction strengths.

physics.chem-ph↗

Block2: a comprehensive open source framework to develop and apply state-of-the-art DMRG algorithms in electronic structure and beyond

Block2 is an open source framework to implement and perform density matrix renormalization group and matrix product state algorithms. Out-of-the-box it supports the eigenstate, time-dependent, response, and finite-temperature algorithms. In addition, it carries special optimizations for ab initio electronic structure Hamiltonians and implements many quantum chemistry extensions to the density matrix renormalization group, such as dynamical correlation theories. The code is designed with an emphasis on flexibility, extensibility, and efficiency, and to support integration with external numerical packages. Here we explain the design principles and currently supported features and present numerical examples in a range of applications.

physics.chem-ph↗

Finite-temperature simulations of strongly correlated systems

This thesis describes several topics related to finite temperature studies of strongly correlated systems: finite temperature density matrix embedding theory (FT-DMET), finite temperature metal-insulator transition, and quantum algorithms including quantum imaginary time evolution (QITE), quantum Lanczos (QLanczos), and quantum minimally entangled typical thermal states (QMETTS) algorithms. While the absolute zero temperature is not reachable, studies of physical and chemical problems at finite temperatures, especially at low temperature, is essential for understanding the quantum behaviors of materials in realistic conditions. Here we define low temperature as the temperature regime where the quantum effect is not largely dissipated due to thermal fluctuation. Treatment of systems at low temperatures is especially difficult compared to both high temperatures - where classical approximation can be applied - and zero temperatures where only the ground state is required to describe the system of interest.

cond-mat.str-el↗

When to Reject a Ground State Preparation Algorithm

In recent years substantial research effort has been devoted to quantum algorithms for ground state energy estimation (GSEE) in chemistry and materials. Given the many heuristic and non-heuristic methods being developed, it is challenging to assess what combination of these will ultimately be used in practice. One important metric for assessing utility is runtime. For most GSEE algorithms, the runtime depends on the ground state preparation (GSP) method. Towards assessing the utility of various combinations of GSEE and GSP methods, we asked under which conditions a GSP method should be accepted over a reference method, such as the Hartree-Fock state. We introduce a criteria for accepting or rejecting a GSP method for the purposes of GSEE. We consider different GSP methods ranging from heuristics to algorithms with provable performance guarantees and perform numerical simulations to benchmark their performance on different chemical systems, starting from small molecules like the hydrogen atom to larger systems like the jellium. In the future this approach may be used to abandon certain VQE ansatzes and other heursitics. Yet so far our findings do not provide evidence against the use of VQE and more expensive heuristic methods, like the low-depth booster. This work sets a foundation from which to further explore the requirements to achieve quantum advantage in quantum chemistry.

quant-ph↗

Variational quantum iterative power algorithms for global optimization

We introduce a family of variational quantum algorithms called quantum iterative power algorithms (QIPA) that outperform existing hybrid near-term quantum algorithms of the same kind. We demonstrate the capabilities of QIPA as applied to three different global-optimization numerical experiments: the ground-state optimization of the $H_2$ molecular dissociation, search of the transmon qubit ground-state, and biprime factorization. Since our algorithm is hybrid, quantum/classical technologies such as error mitigation and adaptive variational ansatzes can easily be incorporated into the algorithm. Due to the shallow quantum circuit requirements, we anticipate large-scale implementation and adoption of the proposed algorithm across current major quantum hardware.

quant-ph↗

SELFIES and the future of molecular string representations

Artificial intelligence (AI) and machine learning (ML) are expanding in popularity for broad applications to challenging tasks in chemistry and materials science. Examples include the prediction of properties, the discovery of new reaction pathways, or the design of new molecules. The machine needs to read and write fluently in a chemical language for each of these tasks. Strings are a common tool to represent molecular graphs, and the most popular molecular string representation, SMILES, has powered cheminformatics since the late 1980s. However, in the context of AI and ML in chemistry, SMILES has several shortcomings -- most pertinently, most combinations of symbols lead to invalid results with no valid chemical interpretation. To overcome this issue, a new language for molecules was introduced in 2020 that guarantees 100\% robustness: SELFIES (SELF-referencIng Embedded Strings). SELFIES has since simplified and enabled numerous new applications in chemistry. In this manuscript, we look to the future and discuss molecular string representations, along with their respective opportunities and challenges. We propose 16 concrete Future Projects for robust molecular representations. These involve the extension toward new chemical domains, exciting questions at the interface of AI and robust languages and interpretability for both humans and machines. We hope that these proposals will inspire several follow-up works exploiting the full potential of molecular string representations for the future of AI in chemistry and materials science.

physics.chem-ph↗

Towards Domain Generalization in Object Detection

Despite the striking performance achieved by modern detectors when training and test data are sampled from the same or similar distribution, the generalization ability of detectors under unknown distribution shifts remains hardly studied. Recently several works discussed the detectors' adaptation ability to a specific target domain which are not readily applicable in real-world applications since detectors may encounter various environments or situations while pre-collecting all of them before training is inconceivable. In this paper, we study the critical problem, domain generalization in object detection (DGOD), where detectors are trained with source domains and evaluated on unknown target domains. To thoroughly evaluate detectors under unknown distribution shifts, we formulate the DGOD problem and propose a comprehensive evaluation benchmark to fill the vacancy. Moreover, we propose a novel method named Region Aware Proposal reweighTing (RAPT) to eliminate dependence within RoI features. Extensive experiments demonstrate that current DG methods fail to address the DGOD problem and our method outperforms other state-of-the-art counterparts.

cs.CV↗

Ground-state phase diagram of the three-band Hubbard model from density matrix embedding theory

We determine the ground-state phase diagram of the three-band Hubbard model across a range of model parameters using density matrix embedding theory. We study the atomic-scale nature of the antiferromagnetic (AFM) and superconducting (SC) orders, explicitly including the oxygen degrees of freedom. All parametrizations of the model display AFM and SC phases, but the decay of AFM order with doping is too slow compared to the experimental phase diagram, and further, coexistence of AFM and SC orders occurs in all parameter sets. The local magnetic moment localizes entirely at the copper sites. The magnetic phase diagram is particularly sensitive to $Δ_{pd}$ and $t_{pp}$, and existing estimates of the charge transfer gap $Δ_{pd}$ appear too large in so-called minimal model parametrizations. The electron-doped side of the phase diagram is qualitatively distinct from hole-doped side and we find an unusual two-peak structure in the SC in the full model parametrization. Examining the SC order at the atomic scale, within the larger scale $d_{x^2 - y^2}$-wave SC pairing order between Cu-Cu and O-O, we also observe a local $p_{x (y)}$ [or $d_{xz (yz)}$]-symmetry modulation of the pair density on the Cu-O bonds. Our work highlights some of the features that arise in a three-band versus one-band picture, the role of the oxygen degrees of freedom in new kinds of atomic-scale SC orders, and the necessity of re-evaluating current parametrizations of the three-band Hubbard model.

cond-mat.str-el↗

Recent developments in the PySCF program package

PYSCF is a Python-based general-purpose electronic structure platform that both supports first-principles simulations of molecules and solids, as well as accelerates the development of new methodology and complex computational workflows. The present paper explains the design and philosophy behind PYSCF that enables it to meet these twin objectives. With several case studies, we show how users can easily implement their own methods using PYSCF as a development environment. We then summarize the capabilities of PYSCF for molecular and solid-state simulations. Finally, we describe the growing ecosystem of projects that use PYSCF across the domains of quantum chemistry, materials science, machine learning and quantum information science.

physics.chem-ph↗

Determining eigenstates and thermal states on a quantum computer using quantum imaginary time evolution

The accurate computation of Hamiltonian ground, excited, and thermal states on quantum computers stands to impact many problems in the physical and computer sciences, from quantum simulation to machine learning. Given the challenges posed in constructing large-scale quantum computers, these tasks should be carried out in a resource-efficient way. In this regard, existing techniques based on phase estimation or variational algorithms display potential disadvantages; phase estimation requires deep circuits with ancillae, that are hard to execute reliably without error correction, while variational algorithms, while flexible with respect to circuit depth, entail additional high-dimensional classical optimization. Here, we introduce the quantum imaginary time evolution and quantum Lanczos algorithms, which are analogues of classical algorithms for finding ground and excited states. Compared to their classical counterparts, they require exponentially less space and time per iteration, and can be implemented without deep circuits and ancillae, or high-dimensional optimization. We furthermore discuss quantum imaginary time evolution as a subroutine to generate Gibbs averages through an analog of minimally entangled typical thermal states. Finally, we demonstrate the potential of these algorithms via an implementation using exact classical emulation as well as through prototype circuits on the Rigetti quantum virtual machine and Aspen-1 quantum processing unit.

quant-ph↗