SearcharxivSearch

arXiv subjects

Furong Xu

Publications and source records attributed to Furong Xu.

At least 19 recordsLinked to original sources

$\textit{Ab Initio}$ Exact Calculation of Strongly Correlated Nucleonic Matter

Dense nucleonic matter is of vital importance for understanding compact stars and inferring the transition into deconfined quark phase. We present $\textit{ab initio}$ exact calculations of infinite nucleonic matter with the state-of-the-art full configuration-interaction quantum Monte Carlo method, enabling us to rigorously benchmark many-body methods and assess the degree to which the nucleonic matter is correlated. Our method has been numerically validated against exact diagonalization within a small model space. Calculations of nucleonic matter using chiral nuclear forces reveal that symmetric nuclear matter is strikingly strongly correlated, raising questions on previous $\textit{ab initio}$ calculations of nuclear matter with many-body expansion truncations and offering insights into simultaneous descriptions of finite nuclei and infinite nucleonic matter from first principles.

nucl-th

Impact of tensor-rank components of chiral three-nucleon forces on the single-particle structure of calcium isotopes

Background: Chiral three-nucleon forces (3NFs) play a key role in the microscopic description of nuclear shell evolution. A recent work introduced an irreducible tensor decomposition of the chiral 3NF at next-to-next-to-leading order and showed that, in $p$-shell nuclei, the enhancement of the $0p_{3/2}$--$0p_{1/2}$ spin--orbit (SO) splitting is mainly driven by its rank-1 component. Purpose: We extend the aforementioned analysis to the $0f1p$ shell to investigate whether the same mechanism persists in a heavier valence space, and how the different tensor-rank components of the 3NF affect structure properties of calcium isotopes. Methods: Effective shell-model Hamiltonians for neutrons outside the doubly magic $^{40}$Ca core are derived from chiral two-nucleon force plus 3NF. The latter is progressively included through its rank-$λ$ components ($λ=0,1,2,3$), allowing us to isolate their impact on the evolution of the neutron single-particle structure. Results: The significant enhancement of the SO splittings for both $1p$ and $0f$ orbitals produced by the chiral 3NF is mainly induced by its rank-1 component. The rank-2 term gives a smaller contribution, while the rank-3 term is negligible. The rank-0 component, and to a lesser extent the rank-1 component, are found to play an important role in determining the spacings between orbitals with different orbital angular momenta. All modifications induced by the 3NF in the single-particle structure have a relevant impact on the shell-closure properties of $^{48}$Ca. Conclusions: The dominance of the rank-1 two-pion-exchange component of the 3NF in explaining the enhancement of SO splitting -- previously identified in the $p$ shell -- persists in the $0f1p$ shell. Observed effects of the 3NF related to the different angular-momentum dependence of the orbitals are shown to arise essentially from their rank-0 and rank-1 components.

nucl-th

Stochastic Similarity Renormalization Group

By integrating the quantum Monte Carlo technique into the similarity renormalization group (SRG), we have developed a stochastic SRG framework (SRGQMC) capable of both free-space two-body and in-medium many-body evolutions. This approach circumvents the combinatorial tensor-space explosion of many-body flow equations by mapping continuous unitary transformations onto an ensemble of signed random walkers. We benchmark the SRGQMC against deterministic free-space SRG evolutions of realistic nucleon-nucleon (NN) interactions, as well as against in-medium SRG (IMSRG) many-body calculations with the Richardson pairing model at two- and three-body levels [IMSRG(2)/(3)]. While a deterministic extension to the four-body level [IMSRG(4)] remains unfeasible due to prohibitive computational costs, we have achieved the first IMSRG(4) calculation by using the stochastic technique, demonstrating a substantial improvement toward the full configuration-interaction limit. This stochastic framework provides a practical pathway to higher-order IMSRG calculations.

nucl-th

Full configuration interaction quantum Monte Carlo for accurate $\textit{ab initio}$ nuclear structure calculations: algorithms and calculation details

Full configuration interaction quantum Monte Carlo (FCIQMC) is a stochastic many-body solver that has been widely applied to electronic, molecular, and condensed-matter systems. In this work we apply FCIQMC to $\textit{ab initio}$ nuclear structure calculations using interactions derived from chiral effective field theory. We describe the algorithm in detail, including imaginary-time propagation, excitation generation, estimator choices, the initiator approximation with adaptive shift correction, and reduced-density-matrix (RDM) sampling. Benchmark calculations in small model spaces, where deterministic full configuration interaction (FCI) results are available, validate the stochastic calculation of energies, radii, and RDM-based pure estimators. For large model spaces, we analyze the residual finite-walker bias through systematic walker-number convergence and infinite-walker extrapolations. We also demonstrate that FCIQMC can be extended beyond ground-state calculations by computing the low-lying spectrum of $^6$Li.

nucl-th

Full Configuration Interaction Quantum Monte Carlo for Accurate $\textit{Ab Initio}$ Nuclear Structure Calculations

We introduce novel full configuration interaction quantum Monte Carlo (FCIQMC) as an accurate many-body solver for $\textit{ab initio}$ nuclear structure calculations. This stochastic approach directly samples the exact wave function in the full configuration space, enabling high-fidelity treatment of high-order many-body correlations in strongly interacting nuclear systems. Using interactions from chiral effective field theory, we have computed ground-state energies and charge radii of $^4$He, $^8$Be, $^{12}$C and $^{16}$O with sub-percent-level many-body uncertainties. These results establish FCIQMC as a stochastic full-configuration-space solver capable of treating systems beyond the reach of the conventional no-core shell model, and as an accurate benchmark for truncated many-body expansion methods.

nucl-th

Chiral three-nucleon forces for the new local position-space two-nucleon potential in $\textit{ab initio}$ many-body calculations

Three-nucleon force (3NF) plays an important role in understanding the structure of finite nuclei and the saturation properties of infinite nuclear matter. More specifically, 3NF should be necessary for each two-nucleon force (2NF) to obtain more accurate description of nuclear systems. 3NF derived from the chiral effective field theory has been successful in $\textit{ab initio}$ calculations of atomic nuclei. Most of established chiral nuclear forces have a nonlocal form in the momentum space. In this work, we construct a companion chiral 3NF specifically tailored to the new Idaho local position-space 2NF, and calculate binding energies and radii of nuclei up to $^{132}$Sn. We find that a chiral 3NF with hybrid local and nonlocal regulators has advantages in improving the nuclear structure calculations of both binding energies and radii with the new Idaho 2NF. The two low-energy constants of 3NF are constrained by the ground-state energies of $^3$H and $^{16}$O as suggested in a recent work.

nucl-th

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria

Aligning multimodal generative models with human preferences demands reward signals that respect the compositional, multi-dimensional structure of human judgment. Prevailing RLHF approaches reduce this structure to scalar or pairwise labels, collapsing nuanced preferences into opaque parametric proxies and exposing vulnerabilities to reward hacking. While recent Rubrics-as-Reward (RaR) methods attempt to recover this structure through explicit criteria, generating rubrics that are simultaneously reliable, scalable, and data-efficient remains an open problem. We introduce Auto-Rubric as Reward (ARR), a framework that reframes reward modeling from implicit weight optimization to explicit, criteria-based decomposition. Before any pairwise comparison, ARR externalizes a VLM's internalized preference knowledge as prompt-specific rubrics, translating holistic intent into independently verifiable quality dimensions. This conversion of implicit preference structure into inspectable, interpretable constraints substantially suppresses evaluation biases including positional bias, enabling both zero-shot deployment and few-shot conditioning on minimal supervision. To extend these gains into generative training, we propose Rubric Policy Optimization (RPO), which distills ARR's structured multi-dimensional evaluation into a robust binary reward, replacing opaque scalar regression with rubric-conditioned preference decisions that stabilize policy gradients. On text-to-image generation and image editing benchmarks, ARR-RPO outperforms pairwise reward models and VLM judges, demonstrating that explicitly externalizing implicit preference knowledge into structured rubrics achieves more reliable, data-efficient multimodal alignment, revealing that the bottleneck is the absence of a factorized interface, not a deficit of knowledge.

cs.AI

Stochastic many-body perturbation theory for high-order calculations

High-order perturbative $\textit{ab initio}$ calculations are challenging due to the rapidly growing configuration space and the difficulty of assessing convergence. In this letter, we introduce perturbation theory quantum Monte Carlo (PTQMC), a stochastic approach designed to compute high-order many-body perturbative corrections. By representing the perturbative wave function with random walkers in configuration space, PTQMC avoids the exponential scaling inherent to conventional constructions of high-rank excitation operators. Benchmark calculations for the Richardson pairing model demonstrate that PTQMC accurately reproduces exact many-body perturbation theory (MBPT) coefficients up to 16th order, even in strongly divergent regimes. We further show that combining PTQMC with series resummation techniques yields stable and precise energy estimates in cases where the straightforward perturbative series fails. Finally, we propose the effective number of configurations, $e^{S}$, as a global measure of perturbative wave-function complexity that can be directly extracted within PTQMC. We demonstrate that the saturation behavior of $e^{S}$ provides a more reliable indicator of the validity of perturbative expansions than energy convergence alone.

nucl-th

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.

cs.CV

Toward $\textit{Ab Initio}$ Quantum Simulations of Atomic Nuclei Using Noisy Qubits

Quantum computers are expected to provide a ultimate solver for quantum many-body systems, although it is a tremendous challenge to achieve that goal on current noisy quantum devices. This work illustrated quantum simulations of ab initio no-core shell model calculations of $^3$H with chiral two-nucleon and three-nucleon forces. The measurement costs are remarkably reduced by using the general commutativity measurement together with the asymptotic optimization. In addition, the noise causes serious contaminations of configurations with undesired particle numbers, and the accuracies are much improved by applying the particle number projected measurement. By tackling the efficiency and noise issues, this work demonstrated a substantial step toward ab initio quantum computing of atomic nuclei.

nucl-th

Complex-energy eigenvector continuation for nuclear many-body broad resonances

Broad resonances are a unique phenomenon in nuclear many-body systems. Theoretical studies usually involve the continuum degree of freedom, which drastically increases the model space of calculations, and may lead to non-convergence or instability of computations. In this paper, we present the extension of the eigenvector continuation (EC) method to the complex-energy space to treat the broad resonances of open quantum systems of nuclei. EC provides an efficient method to predict the solution of a large-space many-body problem within a small subspace. Using only a few bound and narrow resonance solutions as input in EC, we can obtain the solution of a broad resonance. We have applied the complex-energy EC to the broad resonances of $^4$H, four-neutron $^4n$, $^6$He and $^7$He systems.

nucl-th

Ming-Omni: A Unified Multimodal Model for Perception and Generation

We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. Ming-Omni employs dedicated encoders to extract tokens from different modalities, which are then processed by Ling, an MoE architecture equipped with newly proposed modality-specific routers. This design enables a single model to efficiently process and fuse multimodal inputs within a unified framework, thereby facilitating diverse tasks without requiring separate models, task-specific fine-tuning, or structural redesign. Importantly, Ming-Omni extends beyond conventional multimodal models by supporting audio and image generation. This is achieved through the integration of an advanced audio decoder for natural-sounding speech and Ming-Lite-Uni for high-quality image generation, which also allow the model to engage in context-aware chatting, perform text-to-speech conversion, and conduct versatile image editing. Our experimental results showcase Ming-Omni offers a powerful solution for unified perception and generation across all modalities. Notably, our proposed Ming-Omni is the first open-source model we are aware of to match GPT-4o in modality support, and we release all code and model weights to encourage further research and development in the community.

cs.AI

Non-perturbative calculations of nuclear matter using in-medium similarity renormalization group

The non-perturbative {\it ab initio} calculations of infinite nuclear matter using In-Medium Similarity Renormalization Group (IMSRG) method is developed in this work, which enables calculations with chiral two and three-nucleon forces at N$^2$LO and N$^3$LO. Results from the many-body perturbation theory at different orders and coupled-cluster theory are also presented for comparison. It is shown that different many-body approaches lead to obvious discrepancies with a harder nuclear interaction for both pure neutron matter and symmetric nuclear matter. This work provides a novel alternative infrastructure for future studies of dense nuclear matter and strongly-correlated many-body systems.

nucl-th

Ultra-short lifetime isomer studies from photonuclear reactions using laser-driven ultra-intense γ-ray

Isomers, ubiquitous populations of relatively long-lived nuclear excited states, play a crucial role in nuclear physics. However, isomers with half-life times of several seconds or less barely had experimental cross section data due to the lack of a suitable measuring method. We report a method of online γ spectroscopy for ultra-short-lived isomers from photonuclear reactions using laser-driven ultra-intense γ-rays. The fastest time resolution can reach sub-ps level with γ-ray intensities >10^{19}/s ({\geqslant} 8 MeV). The ^{115}In(γ, n)^{114m2}In reaction (T_{1/2} = 43.1 ms) was first measured in the high-energy region which shed light on the nuclear structure studies of In element. Simulations showed it would be an efficient way to study ^{229m}Th (T_{1/2} = 7 μs), which is believed to be the next generation of nuclear clock. This work offered a unique way of gaining insight into ultra-short lifetimes and promised an effective way to fill the gap in relevant experimental data.

nucl-ex

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative scarcity of large-scale pretraining datasets. Toward this end, we introduce a comprehensive bilingual (Chinese-English) dataset BM-6B with over 6 billion image-text pairs, aimed at enhancing multimodal foundation models to well understand images in both languages. To handle such a scale of dataset, we propose a novel grouped aggregation approach for image-text contrastive loss computation, which reduces the communication overhead and GPU memory demands significantly, facilitating a 60% increase in training speed. We pretrain a series of bilingual image-text foundation models with an enhanced fine-grained understanding ability on BM-6B, the resulting models, dubbed as $M^2$-Encoders (pronounced "M-Square"), set new benchmarks in both languages for multimodal retrieval and classification tasks. Notably, Our largest $M^2$-Encoder-10B model has achieved top-1 accuracies of 88.5% on ImageNet and 80.7% on ImageNet-CN under a zero-shot classification setting, surpassing previously reported SoTA methods by 2.2% and 21.1%, respectively. The $M^2$-Encoder series represents one of the most comprehensive bilingual image-text foundation models to date, so we are making it available to the research community for further exploration and development.

cs.CV

Properties of chiral nucleon-nucleon interaction at N$^3$LO with high cutoffs studied by local projection

The chiral nucleon-nucleon ($NN$) interaction at high cutoffs has been plagued by the presence of spurious bound states. In this work, the chiral $NN$ interaction at N$^3$LO is studied by the local projection method as the cutoff increases. The evolution of short-range behaviors of pion-exchange interactions and contact interactions is intuitively demonstrated. The $P$-channel potentials toward high cutoffs appear to be erratic at short ranges to compromise with phase shifts, while such erratic behaviors can be avoided in $S$ and $D$ channels. Furthermore, a chiral $NN$ interaction at N$^3$LO is studied at a cutoff of 700 MeV. The properties of deuteron and triton are testified with this interaction. Such a hard interaction is expected to provide an alternative choice for studies of short-range correlations and high density nuclear matter.

nucl-th

SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impressive results. CLIP imposes a bidirectional constraints on global representation of entire images and sentences. Although IC conducts an unidirectional image-to-text generation on local representation, it lacks any constraint on local text-to-image reconstruction, which limits the ability to understand images at a fine-grained level when aligned with texts. To achieve multimodal alignment from both global and local perspectives, this paper proposes Symmetrizing Contrastive Captioners (SyCoCa), which introduces bidirectional interactions on images and texts across the global and local representation levels. Specifically, we expand a Text-Guided Masked Image Modeling (TG-MIM) head based on ITC and IC heads. The improved SyCoCa can further leverage textual cues to reconstruct contextual images and visual cues to predict textual contents. When implementing bidirectional local interactions, the local contents of images tend to be cluttered or unrelated to their textual descriptions. Thus, we employ an attentive masking strategy to select effective image patches for interaction. Extensive experiments on five vision-language tasks, including image-text retrieval, image-captioning, visual question answering, and zero-shot/finetuned image classification, validate the effectiveness of our proposed method.

cs.CV

Text as Image: Learning Transferable Adapter for Multi-Label Classification

Pre-trained vision-language models have notably accelerated progress of open-world concept recognition. Their impressive zero-shot ability has recently been transferred to multi-label image classification via prompt tuning, enabling to discover novel labels in an open-vocabulary manner. However, this paradigm suffers from non-trivial training costs, and becomes computationally prohibitive for a large number of candidate labels. To address this issue, we note that vision-language pre-training aligns images and texts in a unified embedding space, making it potential for an adapter network to identify labels in visual modality while be trained in text modality. To enhance such cross-modal transfer ability, a simple yet effective method termed random perturbation is proposed, which enables the adapter to search for potential visual embeddings by perturbing text embeddings with noise during training, resulting in better performance in visual modality. Furthermore, we introduce an effective approach to employ large language models for multi-label instruction-following text generation. In this way, a fully automated pipeline for visual label recognition is developed without relying on any manual data. Extensive experiments on public benchmarks show the superiority of our method in various multi-label classification tasks.

cs.CV