SearcharxivSearch

arXiv subjects

Fanjie Xu

Publications and source records attributed to Fanjie Xu.

15 recordsLinked to original sources

Uni-XAS: Alignment-Driven Bidirectional Multimodal Learning for X-ray Absorption Spectroscopy

X-ray absorption spectroscopy (XAS) is a key technique for probing local atomic environments, yet learning based modeling must bridge two heterogeneous modalities: 1D continuous spectra and 3D atomic structures. Existing approaches typically decouple forward spectrum prediction and inverse structure inference into separate regression tasks, hindering shared representation learning. Moreover, severe permutation ambiguity among identical atoms often limits inverse modeling to coarse structure descriptors rather than explicit 3D structure generation. In this work, we present Uni-XAS, a unified benchmark and learning framework that reframes bidirectional XAS modeling as a cross-modal alignment and conditional generation problem. We first propose XASLip, an alignment recipe coupling a physics-aware spectral encoder with an absorberaware manifold optimization strategy to resolve fine-grained intra-element coordination variations. Building upon this shared latent space, we formulate forward prediction as anchored absolute-spectrum generation via retrieval-augmented decoding, effectively preventing physical scale collapse and energy drift. For the inherently ill-posed inverse problem, we introduce Permutation-Rectified Flow Matching, which integrates type-wise optimal transport into a continuous generative flow to provide a principled solution to ligand permutation ambiguity without relying on heavy high-order equivariant architectures. Evaluated on a largescale standardized benchmark of 328,839 structure-spectrum pairs, Uni-XAS demonstrates strong performance in cross-modal retrieval, accurate absolute-spectrum prediction, and composition-conditional 3D structure generation, establishing a scalable, reproducible, and protocol-consistent foundation for multimodal learning and standardized evaluation in scientific spectroscopy.

cond-mat.mtrl-sci

Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent

Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown molecules remains a bottleneck reliant on human expertise. While artificial intelligence has advanced this field, current methods face a critical trade-off: database retrieval cannot identify novel scaffolds, while de novo molecular structure elucidation models operate as black boxes, lacking the atom-level interpretability required for rigorous scientific validation. Here, we present NMRAgent, an evidential reasoning agent powered by large language models (LLMs) that bridges this gap by integrating specialized spectral analysis tools with chemical knowledge graphs. Unlike previous approaches, NMRAgent mimics the deductive reasoning of human experts: it takes experimental NMR spectra and molecular formula as input, plans the elucidation process, proposes candidate structures, verifies peak-atom consistency, and refines misaligned substructure through formula-aware fragment optimization. Enabled by its evidential reasoning, NMRAgent outperforms state-of-the-art methods, improving top-1 accuracy by 46.5% and Tanimoto similarity by 0.502 on a scaffold-split benchmark with novel scaffolds in the test set. Besides, we demonstrate the agent's practical utility by elucidating the structures of two previously unknown natural products isolated from Hydrangea davidii and Vitex trifolia, and by correcting structural misassignments in established literature. By combining high-accuracy prediction with transparent and evidence-based reasoning, NMRAgent establishes a new paradigm for interpretable AI in analytical chemistry.

cs.LG

A large-scale foundation model enables simulation-to-real adaptation for nuclear magnetic resonance-based molecular structure analysis

Nuclear Magnetic Resonance (NMR) spectroscopy is a powerful tool for molecular structure analysis, and spectral artificial intelligence offers great potential for its rapid and automated interpretation. However, the scarcity of experimental NMR datasets has constrained deep learning in this domain to narrow, task-specific applications that lack broad generalization. Here, we introduce UltraNMR, a large-scale foundation model for NMR that leverages the intrinsic properties of NMR spectra to learn generalizable spectral representations. We collected 158 million paired simulated $^{1}$H and $^{13}$C NMR spectra to train UltraNMR, employing multiple domain-specific pre-training objectives. UltraNMR captures both intra-spectral and inter-spectral dependencies, enabling seamless simulation-to-real adaptation. We demonstrate that adapting UltraNMR to a range of molecular structure analysis tasks on experimental NMR spectra consistently yields state-of-the-art performance and clearly outperforms UltraNMR variants trained directly on downstream data without simulation pre-training. We also construct a large-scale NMR spectral vector library by encoding simulated NMR spectra using UltraNMR, covering 94 million unique molecules and enabling effective structure-aware retrieval. In real-world applications, UltraNMR facilitates the structural elucidation of two previously unknown natural products from Chinese herbal medicines recorded in the Chinese Pharmacopoeia. These results suggest that large-scale simulation pre-training can effectively bridge the simulation-to-real gap, enabling robust and generalizable molecular structure analysis of real-world NMR spectra.

physics.chem-ph

Solving the inverse problem of X-ray absorption spectroscopy via physics-informed deep learning

Resolving transient atomic configurations in non-crystalline or dynamic environments remains a fundamental bottleneck in the physical sciences. While X-ray absorption spectroscopy (XAS) is a premier probe of local structure, inverting spectra into structural descriptors is a notoriously ill-posed problem due to inherent many-to-one mapping. Here, we present the Spectral Pattern Translator (SPT), a physics-informed deep learning framework that establishes a robust bridge between large-scale theoretical datasets and experimental reality. Our strategy exploits the Fourier duality between spectral energy oscillations and spatial scattering paths to overcome the "simulation-to-experiment" gap. By decomposing spectra into frequency domains, SPT effectively isolates robust structural coordination signals from the destabilizing noise inherent in experimental data. Trained on a massive library of diverse atomic environments, this approach achieves state-of-the-art accuracy in resolving continuous phase transitions in battery cathodes and deciphering local order in amorphous materials. With millisecond-scale latency, SPT removes the primary computational barrier to autonomous materials discovery, establishing a robust, noise-resilient engine for closed-loop robotic chemistry.

cond-mat.mtrl-sci

SpecXMaster Technical Report

Intelligent spectroscopy serves as a pivotal element in AI-driven closed-loop scientific discovery, functioning as the critical bridge between matter structure and artificial intelligence. However, conventional expert-dependent spectral interpretation encounters substantial hurdles, including susceptibility to human bias and error, dependence on limited specialized expertise, and variability across interpreters. To address these challenges, we propose SpecXMaster, an intelligent framework leveraging Agentic Reinforcement Learning (RL) for NMR molecular spectral interpretation. SpecXMaster enables automated extraction of multiplicity information from both 1H and 13C spectra directly from raw FID (free induction decay) data. This end-to-end pipeline enables fully automated interpretation of NMR spectra into chemical structures. It demonstrates superior performance across multiple public NMR interpretation benchmarks and has been refined through iterative evaluations by professional chemical spectroscopists. We believe that SpecXMaster, as a novel methodological paradigm for spectral interpretation, will have a profound impact on the organic chemistry community.

cs.LG

Experimental Powder X-ray Diffraction Crystal Structure Determination with RealPXRD-Solver

Determining crystal structures from experimental powder X-ray diffraction data remains challenging because peak overlap, preferred orientation, and impurity phases obscure atomic arrangements. We present RealPXRD-Solver, a generative model trained on 6,250,238 theoretical structures with experiment-mimicking augmentations and a universal encoder of d-spacing--intensity fingerprints, enabling both lattice-conditioned and lattice-free inference. RealPXRD-Solver reaches a 98.3% Top-20 match rate on a 10,000-structure theoretical benchmark and achieves Top-1/Top-20 accuracies of 77.9%/91.9% on CNRS and 78.8%/92.9% on RRUFF experimental datasets, and it solved 39 previously unreported Powder Diffraction File entries.

cond-mat.mtrl-sci

NMRPeak: a ready-to-use intelligent system for molecular structure elucidation enabled by synergistic cross-modal learning

One-dimensional nuclear magnetic resonance (NMR) spectroscopy is essential for molecular structure elucidation in organic synthesis, drug discovery, natural product characterization, and metabolomics, yet its interpretation remains heavily dependent on expert knowledge and difficult to scale. Although machine learning has been applied to NMR spectrum prediction, library retrieval, and structure generation, these tasks have evolved in isolation using simulated data and incompatible spectral representations, limiting their utility under real experimental scenarios. Here we present NMRPeak, a unified cross-modal learning system that integrates these three tasks through experimentally grounded design. We curate approximately 1.8 million experimental and simulated spectra to construct the largest benchmark for NMR-based structure elucidation and systematically quantify the distribution shift between these domains. We introduce a chemically-aware adaptive tokenizer that dynamically balances discretization granularity to preserve spectral semantics while controlling vocabulary size, and an assignment-free peak-aware similarity metric that enables direct comparison between predicted and experimental spectra. Through a unified molecule-to-spectrum paradigm and synergistic coupling of prediction, retrieval, and generation modules, NMRPeak achieves transformative performance on experimental benchmarks: it overcomes the longstanding simulation-to-experiment gap in spectrum prediction while delivering over 95% top-1 accuracy in molecular retrieval and approximately 75% top-1 accuracy in stereochemistry-aware de novo structure generation. These capabilities establish a foundation for automated, high-throughput molecular structure elucidation in organic synthesis, drug discovery, and chemical biology.

cond-mat.mtrl-sci

NMR-Solver: Automated Structure Elucidation via Large-Scale Spectral Matching and Physics-Guided Fragment Optimization

Nuclear Magnetic Resonance (NMR) spectroscopy is one of the most powerful and widely used tools for molecular structure elucidation in organic chemistry. However, the interpretation of NMR spectra to determine unknown molecular structures remains a labor-intensive and expertise-dependent process, particularly for complex or novel compounds. Although recent methods have been proposed for molecular structure elucidation, they often underperform in real-world applications due to inherent algorithmic limitations and limited high-quality data. Here, we present NMR-Solver, a practical and interpretable framework for the automated determination of small organic molecule structures from $^1$H and $^{13}$C NMR spectra. Our method introduces an automated framework for molecular structure elucidation, integrating large-scale spectral matching with physics-guided fragment-based optimization that exploits atomic-level structure-spectrum relationships in NMR. We evaluate NMR-Solver on simulated benchmarks, curated experimental data from the literature, and real-world experiments, demonstrating its strong generalization, robustness, and practical utility in challenging, real-life scenarios. NMR-Solver unifies computational NMR analysis, deep learning, and interpretable chemical reasoning into a coherent system. By incorporating the physical principles of NMR into molecular optimization, it enables scalable, automated, and chemically meaningful molecular identification, establishing a generalizable paradigm for solving inverse problems in molecular science.

physics.chem-ph

On an extension of Shlyk's theorem

In this paper, we prove that the intersection of all non-nilpotent maximal subgroups of a non-solvable group containing the normalizer of some Sylow subgroup is nilpotent, which provides an extension of Shlyk's theorem.

math.GR

Finite groups in which some particular non-nilpotent maximal invariant subgroups have indices a prime-power

Let $A$ and $G$ be finite groups such that $A$ acts coprimely on $G$ by automorphisms, assume that $G$ has a maximal $A$-invariant subgroup $M$ that is a direct product of some isomorphic simple groups, we prove that if $G$ has a non-trivial $A$-invariant normal subgroup $N$ such that $N\leq M$ and every non-nilpotent maximal $A$-invariant subgroup $K$ of $G$ not containing $N$ has index a prime-power and the projective special linear group $PSL_2(7)$ is not a composition factor of $G$, then $G$ is solvable.

math.GR

Towards a Unified Benchmark and Framework for Deep Learning-Based Prediction of Nuclear Magnetic Resonance Chemical Shifts

The study of structure-spectrum relationships is essential for spectral interpretation, impacting structural elucidation and material design. Predicting spectra from molecular structures is challenging due to their complex relationships. Herein, we introduce NMRNet, a deep learning framework using the SE(3) Transformer for atomic environment modeling, following a pre-training and fine-tuning paradigm. To support the evaluation of NMR chemical shift prediction models, we have established a comprehensive benchmark based on previous research and databases, covering diverse chemical systems. Applying NMRNet to these benchmark datasets, we achieve state-of-the-art performance in both liquid-state and solid-state NMR datasets, demonstrating its robustness and practical utility in real-world scenarios. This marks the first integration of solid and liquid state NMR within a unified model architecture, highlighting the need for domainspecific handling of different atomic environments. Our work sets a new standard for NMR prediction, advancing deep learning applications in analytical and structural chemistry.

physics.comp-ph

Finite groups with some particular maximal invariant subgroups being nilpotent or all non-nilpotent maximal invariant subgroups being normal

Let $A$ and $G$ be finite groups such that $A$ acts coprimely on $G$ by automorphisms. We provide a complete classification of a finite group $G$ in which every maximal $A$-invariant subgroup containing the normalizer of some $A$-invariant Sylow subgroup is nilpotent. Moreover, we show that both the hypothesis that every maximal $A$-invariant subgroup of $G$ containing the normalizer of some $A$-invariant Sylow subgroup is nilpotent and the hypothesis that every non-nilpotent maximal $A$-invariant subgroup of $G$ is normal are equivalent.

math.GR

On generalizations of Iwasawa's theorem

Iwasawa's theorem indicates that a finite group $G$ is supersolvable if and only if all maximal chains of the identity in $G$ have the same length. As generalizations of Iwasawa's theorem, we provide some characterizations of the structure of a finite group $G$ in which all maximal chains of every minimal subgroup have the same length. Moreover, let $δ(G)$ be the number of subgroups of $G$ all of whose maximal chains in $G$ do not have the same length, we prove that $G$ is a non-solvable group with $δ(G)\leq 16$ if and only if $G\cong A_5$.

math.GR

End-to-End Crystal Structure Prediction from Powder X-Ray Diffraction

Powder X-ray diffraction (PXRD) is a prevalent technique in materials characterization. While the analysis of PXRD often requires extensive human manual intervention, and most automated method only achieved at coarse-grained level. The more difficult and important task of fine-grained crystal structure prediction from PXRD remains unaddressed. This study introduces XtalNet, the first equivariant deep generative model for end-to-end crystal structure prediction from PXRD. Unlike previous crystal structure prediction methods that rely solely on composition, XtalNet leverages PXRD as an additional condition, eliminating ambiguity and enabling the generation of complex organic structures with up to 400 atoms in the unit cell. XtalNet comprises two modules: a Contrastive PXRD-Crystal Pretraining (CPCP) module that aligns PXRD space with crystal structure space, and a Conditional Crystal Structure Generation (CCSG) module that generates candidate crystal structures conditioned on PXRD patterns. Evaluation on two MOF datasets (hMOF-100 and hMOF-400) demonstrates XtalNet's effectiveness. XtalNet achieves a top-10 Match Rate of 90.2% and 79% for hMOF-100 and hMOF-400 in conditional crystal structure prediction task, respectively. XtalNet enables the direct prediction of crystal structures from experimental measurements, eliminating the need for manual intervention and external databases. This opens up new possibilities for automated crystal structure determination and the accelerated discovery of novel materials.

physics.chem-ph