SearcharxivSearch

arXiv subjects

Xiang Jiang

Publications and source records attributed to Xiang Jiang.

At least 19 recordsLinked to original sources

Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture Generation

Moonshine is an autonomous agent whose central objective is to generate mathematical conjectures. Its core capability is to extract structure from classical problems, distill new concepts, and formulate conjectures of mathematical significance. Rather than treating the solution of a single proposition as its endpoint, Moonshine builds an extensible theoretical framework through conjecture generation, bridge building, and obstacle identification. This article uses Moonshine's exploration of the Jacobian conjecture as an example. It shows how the central logic of whether local nondegeneracy can force global injectivity is transferred to one-hidden-layer affine-ridge sigmoid networks. This leads to the formulation of the \emph{Neural Jacobian Conjecture} (NJC): if such a network has strictly positive Jacobian determinant on the whole space, then it must be globally injective. By invoking GPT-5.5-pro and DeepSeek-V4-pro separately, Moonshine obtained independent complete proofs for the case \(N=n+1\). In addition, with the assistance of ChatGPT through interactive use of its web interface with GPT-5.5-pro, a geometric-topological proof was developed. These results provide preliminary evidence for the plausibility of the conjecture. The general higher-width case \(N\ge n+2\), however, remains unresolved and is left for further investigation. This work illustrates Moonshine's ability to autonomously generate meaningful mathematical problems and make rigorous progress on them.

cs.AI

Can LLM generate interesting mathematical research problems?

This paper is the second one in a series of work on the mathematical creativity of LLM. In the first paper, the authors proposed three criteria for evaluating the mathematical creativity of LLM and constructed a benchmark dataset to measure it. This paper further explores the mathematical creativity of LLM, with a focus on investigating whether LLM can generate valuable and cutting-edge mathematical research problems. We develop an agent to generate unknown problems and produced 665 research problems in differential geometry. Through human verification, we find that many of these mathematical problems are unknown to experts and possess unique research value.

cs.AI

Shell effects in nuclear charge radii based on Skyrme density functionals

A unified description of the charge radii throughout the entire nuclide chart plays an essential role for our understanding of nuclear structure and fundamental nuclear interactions. In this work, the influence of new term, which catches the spirit of neutron and proton pairs condensation around Fermi surface, on the charge radii has been investigated based on the Skyrme density functionals with the effective forces SLy5 and SkM$^{*}$. The differential charge radii of even-even Ca, Ni, Sn, and Pb isotopes are employed to evaluate the validity of this theoretical model. Meanwhile, the results obtained by the relativistic density functional with the effective Lagrangian NL3 are also shown for the quantitative comparison. The calculated results suggest that the modified model can improve the trend of changes of the differential charge radii along Ca, Ni, Sn, and Pb isotopic chains, especially the shell closure effect at the neutron numbers $N=28$, 82 and 126. The shell quenching phenomena of charge radii can also be predicted at the neutron number $N=50$ along the corresponding Ni and Sn isotopes, respectively. The inverted parabolic-like shapes between the two fully filled shells can also be observed, but the amplitude is gradually weakened from Ca to Pb isotopic chains. Combining the existing literatures, it suggests that the discontinuous behavior in nuclear charge radii can be described well by considering the influence of neutron Cooper pairs condensation around Fermi surface.

nucl-th

DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models

To advance the mathematical proficiency of large language models (LLMs), the DeepMath team has launched an open-source initiative aimed at developing an open mathematical LLM and systematically evaluating its mathematical creativity. This paper represents the initial contribution of this initiative. While recent developments in mathematical LLMs have predominantly emphasized reasoning skills, as evidenced by benchmarks on elementary to undergraduate-level mathematical tasks, the creative capabilities of these models have received comparatively little attention, and evaluation datasets remain scarce. To address this gap, we propose an evaluation criteria for mathematical creativity and introduce DeepMath-Creative, a novel, high-quality benchmark comprising constructive problems across algebra, geometry, analysis, and other domains. We conduct a systematic evaluation of mainstream LLMs' creative problem-solving abilities using this dataset. Experimental results show that even under lenient scoring criteria -- emphasizing core solution components and disregarding minor inaccuracies, such as small logical gaps, incomplete justifications, or redundant explanations -- the best-performing model, O3 Mini, achieves merely 70% accuracy, primarily on basic undergraduate-level constructive tasks. Performance declines sharply on more complex problems, with models failing to provide substantive strategies for open problems. These findings suggest that, although current LLMs display a degree of constructive proficiency on familiar and lower-difficulty problems, such performance is likely attributable to the recombination of memorized patterns rather than authentic creative insight or novel synthesis.

cs.AI

NeoQA: Evidence-based Question Answering with Generated News Events

Evaluating Retrieval-Augmented Generation (RAG) in large language models (LLMs) is challenging because benchmarks can quickly become stale. Questions initially requiring retrieval may become answerable from pretraining knowledge as newer models incorporate more recent information during pretraining, making it difficult to distinguish evidence-based reasoning from recall. We introduce NeoQA (News Events for Out-of-training Question Answering), a benchmark designed to address this issue. To construct NeoQA, we generated timelines and knowledge bases of fictional news events and entities along with news articles and Q\&A pairs to prevent LLMs from leveraging pretraining knowledge, ensuring that no prior evidence exists in their training data. We propose our dataset as a new platform for evaluating evidence-based question answering, as it requires LLMs to generate responses exclusively from retrieved evidence and only when sufficient evidence is available. NeoQA enables controlled evaluation across various evidence scenarios, including cases with missing or misleading details. Our findings indicate that LLMs struggle to distinguish subtle mismatches between questions and evidence, and suffer from short-cut reasoning when key information required to answer a question is missing from the evidence, underscoring key limitations in evidence-based reasoning.

cs.CL

Implication of shell quenching in scandium isotopes around N=20

Shell closure structures are commonly observed phenomena associated with nuclear charge radii throughout the nuclide chart. Inspired by recent studies demonstrating that the abrupt change can be clearly observed in the charge radii of the scandium isotopic chain across the neutron number $N=20$, we further review the underlying mechanism of the enlarged charge radii for $^{42}$Sc based on the covariant density functional theory. The pairing correlations are tackled by solving the state-dependent Bardeen-Cooper-Schrieffer equations. Meanwhile, the neutron-proton correlation around the Fermi surface derived from the simultaneously unpaired proton and neutron is appropriately considered in describing the systematic evolution of nuclear charge radii. The calculated results suggest that the abrupt increase in charge radii across the $N=20$ shell closure seems to be improved along the scandium isotopic chain if the strong neutron-proton correlation is properly included.

nucl-th

Fuzzy Cluster-Aware Contrastive Clustering for Time Series

The rapid growth of unlabeled time series data, driven by the Internet of Things (IoT), poses significant challenges in uncovering underlying patterns. Traditional unsupervised clustering methods often fail to capture the complex nature of time series data. Recent deep learning-based clustering approaches, while effective, struggle with insufficient representation learning and the integration of clustering objectives. To address these issues, we propose a fuzzy cluster-aware contrastive clustering framework (FCACC) that jointly optimizes representation learning and clustering. Our approach introduces a novel three-view data augmentation strategy to enhance feature extraction by leveraging various characteristics of time series data. Additionally, we propose a cluster-aware hard negative sample generation mechanism that dynamically constructs high-quality negative samples using clustering structure information, thereby improving the model's discriminative ability. By leveraging fuzzy clustering, FCACC dynamically generates cluster structures to guide the contrastive learning process, resulting in more accurate clustering. Extensive experiments on 40 benchmark datasets show that FCACC outperforms the selected baseline methods (eight in total), providing an effective solution for unsupervised time series learning.

cs.LG

Implication of odd-even staggering in the charge radii of calcium isotopes

Inspired by the profoundly observed odd-even staggering and the inverted parabolic-like shape in charge radii along calcium isotopic chain, the ground state properties of calcium isotopes are investigated by constraining the root-mean-square (rms) charge radii under the covariant energy density functionals with effective forces NL3 and PK1. In this work, the pairing correlations are tackled by solving the state-dependent Bardeen-Cooper-Schrieffer equations. The calculated results suggest that the binding energies obtained by the radius constraint method have been slightly changed by about $0.2\%$. But for charge radii, the corresponding results deriving from NL3 and PK1 forces have been increased by about $1.0\%$ and $2.0\%$, respectively. This means that charge radius is a more sensitive quantity in the calibrated protocol. Meanwhile, it is found that the reproduced charge radii of calcium isotopes are attributed to the rather strong isospin dependence of effective potential. The odd-even oscillation behavior can also be presented in the proton Fermi energies along calcium isotopic family, but keep opposite trends with respect to the corresponding binding energies and charge radii. As encountered in charge radii, the weakened odd-even oscillation behavior is still emerged from the proton Fermi energies at the neutron numbers $N=20$ and $28$ as well, but not in binding energies.

nucl-th

Improved description of nuclear charge radii: Global trends beyond $N=28$ shell closure

Charge radii measured with high accuracy provide a stringent benchmark for characterizing nuclear structure phenomena. In this work, the systematic evolution of charge radii for nuclei with $Z=19$-$29$ is investigated through relativistic mean field theory with effective forces NL3, PK1, and NL3$^{*}$. The neutron-proton ($np$) correlation around Fermi surface originated from the unpaired neutron and proton has been taken into account tentatively in order to reduce the overestimated odd-even staggering of charge radii. This improved method can give an available description of charge radii across $N=28$ shell closure. A remarkable observation is that the charge radii beyond $N=28$ shell closure follow the similarly steep increasing trend, namely irrespective of the number of protons in the nucleus. Especially, the latest results of charge radii for nickel and copper isotopes can be reproduced remarkably well. Along $N=28$ isotonic chain, the sudden increase of charge radii is weakened across $Z=20$, but presented evidently across $Z=28$ closed shell. The abrupt changes of charge radii across $Z=22$ are also shown along $N=32$ and $34$ isotones, but the latter with a less slope. This seems to provide a sensitive indicator to identify the new magicity of a nucleus with universal trend of charge radii.

nucl-th

Evolution of nuclear charge radii in copper and indium isotopes

Systematic trends in nuclear charge radii are of great interest due to universal shell effects and odd-even staggering (OES). The modified root mean square (rms) charge radius formula, which phenomenologically accounts for the formation of neutron-proton ($np$) correlations, is here applied for the first time to the study of odd-$Z$ copper and indium isotopes. Theoretical results obtained by the relativistic mean field (RMF) model with NL3, PK1 and NL3$^{*}$ parameter sets are compared with experimental data. Our results show that both OES and the abrupt changes across $N=50$ and $82$ shell closures are clearly reproduced in nuclear charge radii. The inverted parabolic-like behaviors of rms charge radii can also be described remarkably well between two neutron magic numbers, namely $N=28$ to $50$ for copper isotopes and $N=50$ to $82$ for indium isotopes. This implies that the $np$-correlations play an indispensable role in quantitatively determining the fine structures of nuclear charge radii along odd-$Z$ isotopic chains. Also, our conclusions have almost no dependence on the effective forces.

nucl-th

Odd-even staggering and shell effects of charge radii for nuclei with even $Z$ from $36$ to $38$ and from $52$ to $62$

A unified theoretical model reproducing charge radii of known atomic nuclei plays an essential role in making extrapolations for unknown nuclei. Recently developed new ansatz which phenomenologically takes into account the neutron-proton short-range correlations ($np$-SRCs) can describe the discontinuity properties and odd-even staggering (OES) effect of charge radii along isotopic chains remarkably well. In this work, we further review the modified root-mean-square (rms) charge radii formula in the framework of relativistic mean field (RMF) theory. The charge radii are calculated along various isotopic chains that include the nuclei featuring the $N=50$ and $82$ magic shells. Our results suggest that RMF with and without considering a correction term give an almost similar trend of nuclear size for some isotopic chains with open proton shell, especially the abrupt increases across the strong neutron closed shells and the OES behaviors. This reflects that the $np$-SRCs have almost no influence for some nuclei due to the strong coupling between different levels around Fermi surface. The weakening OES behavior of nuclear charge radii is observed generally at completely filled neutron shells and this may be proposed as a signature of magic indicator.

nucl-th

Hypothesis Disparity Regularized Mutual Information Maximization

We propose a hypothesis disparity regularized mutual information maximization~(HDMI) approach to tackle unsupervised hypothesis transfer -- as an effort towards unifying hypothesis transfer learning (HTL) and unsupervised domain adaptation (UDA) -- where the knowledge from a source domain is transferred solely through hypotheses and adapted to the target domain in an unsupervised manner. In contrast to the prevalent HTL and UDA approaches that typically use a single hypothesis, HDMI employs multiple hypotheses to leverage the underlying distributions of the source and target hypotheses. To better utilize the crucial relationship among different hypotheses -- as opposed to unconstrained optimization of each hypothesis independently -- while adapting to the unlabeled target domain through mutual information maximization, HDMI incorporates a hypothesis disparity regularization that coordinates the target hypotheses jointly learn better target representations while preserving more transferable source knowledge with better-calibrated prediction uncertainty. HDMI achieves state-of-the-art adaptation performance on benchmark datasets for UDA in the context of HTL, without the need to access the source data during the adaptation.

cs.LG

Implicit Class-Conditioned Domain Alignment for Unsupervised Domain Adaptation

We present an approach for unsupervised domain adaptation---with a strong focus on practical considerations of within-domain class imbalance and between-domain class distribution shift---from a class-conditioned domain alignment perspective. Current methods for class-conditioned domain alignment aim to explicitly minimize a loss function based on pseudo-label estimations of the target domain. However, these methods suffer from pseudo-label bias in the form of error accumulation. We propose a method that removes the need for explicit optimization of model parameters from pseudo-labels directly. Instead, we present a sampling-based implicit alignment approach, where the sample selection procedure is implicitly guided by the pseudo-labels. Theoretical analysis reveals the existence of a domain-discriminator shortcut in misaligned classes, which is addressed by the proposed implicit alignment approach to facilitate domain-adversarial learning. Empirical results and ablation studies confirm the effectiveness of the proposed approach, especially in the presence of within-domain class imbalance and between-domain class distribution shift.

cs.LG

Continuous Domain Adaptation with Variational Domain-Agnostic Feature Replay

Learning in non-stationary environments is one of the biggest challenges in machine learning. Non-stationarity can be caused by either task drift, i.e., the drift in the conditional distribution of labels given the input data, or the domain drift, i.e., the drift in the marginal distribution of the input data. This paper aims to tackle this challenge in the context of continuous domain adaptation, where the model is required to learn new tasks adapted to new domains in a non-stationary environment while maintaining previously learned knowledge. To deal with both drifts, we propose variational domain-agnostic feature replay, an approach that is composed of three components: an inference module that filters the input data into domain-agnostic representations, a generative module that facilitates knowledge transfer, and a solver module that applies the filtered and transferable knowledge to solve the queries. We address the two fundamental scenarios in continuous domain adaptation, demonstrating the effectiveness of our proposed approach for practical usage.

cs.LG

Production mechanism of neutron-rich nuclei around $N=126$ in multi-nucleon transfer reaction $^{132}$Sn + $^{208}$Pb

Time-dependent Hartree-Fock approach in three dimensions is employed to study the multi-nucleon transfer reaction $^{132}$Sn + $^{208}$Pb at various incident energies above the Coulomb barrier. Probabilities for different transfer channels are calculated by using particle-number projection method. The results indicate that neutron stripping (transfer from the projectile to the target) and proton pick-up (transfer from the target to the projectile) are favored. Deexcitation of the primary fragments are treated by using the state-of-art statistical code GEMINI++. Primary and final production cross sections of the target-like fragments (with $Z=77$ to $Z=87$) are investigated. The results reveal that fission decay of heavy nuclei plays an important role in the deexcitation process of nuclei with $Z>82$. It is also found that the final production cross sections of neutron-rich (deficient) nuclei slightly (strongly) depend on the incident energy.

nucl-th

On the Importance of Attention in Meta-Learning for Few-Shot Text Classification

Current deep learning based text classification methods are limited by their ability to achieve fast learning and generalization when the data is scarce. We address this problem by integrating a meta-learning procedure that uses the knowledge learned across many tasks as an inductive bias towards better natural language understanding. Based on the Model-Agnostic Meta-Learning framework (MAML), we introduce the Attentive Task-Agnostic Meta-Learning (ATAML) algorithm for text classification. The essential difference between MAML and ATAML is in the separation of task-agnostic representation learning and task-specific attentive adaptation. The proposed ATAML is designed to encourage task-agnostic representation learning by way of task-agnostic parameterization and facilitate task-specific adaptation via attention mechanisms. We provide evidence to show that the attention mechanism in ATAML has a synergistic effect on learning performance. In comparisons with models trained from random initialization, pretrained models and meta trained MAML, our proposed ATAML method generalizes better on single-label and multi-label classification tasks in miniRCV1 and miniReuters-21578 datasets.

cs.LG

TrajectoryNet: An Embedded GPS Trajectory Representation for Point-based Classification Using Recurrent Neural Networks

Understanding and discovering knowledge from GPS (Global Positioning System) traces of human activities is an essential topic in mobility-based urban computing. We propose TrajectoryNet-a neural network architecture for point-based trajectory classification to infer real world human transportation modes from GPS traces. To overcome the challenge of capturing the underlying latent factors in the low-dimensional and heterogeneous feature space imposed by GPS data, we develop a novel representation that embeds the original feature space into another space that can be understood as a form of basis expansion. We also enrich the feature space via segment-based information and use Maxout activations to improve the predictive power of Recurrent Neural Networks (RNNs). We achieve over 98% classification accuracy when detecting four types of transportation modes, outperforming existing models without additional sensory data or location-based prior knowledge.

cs.CV

Camera Fingerprint: A New Perspective for Identifying User's Identity

Identifying user's identity is a key problem in many data mining applications, such as product recommendation, customized content delivery and criminal identification. Given a set of accounts from the same or different social network platforms, user identification attempts to identify all accounts belonging to the same person. A commonly used solution is to build the relationship among different accounts by exploring their collective patterns, e.g., user profile, writing style, similar comments. However, this kind of method doesn't work well in many practical scenarios, since the information posted explicitly by users may be false due to various reasons. In this paper, we re-inspect the user identification problem from a novel perspective, i.e., identifying user's identity by matching his/her cameras. The underlying assumption is that multiple accounts belonging to the same person contain the same or similar camera fingerprint information. The proposed framework, called User Camera Identification (UCI), is based on camera fingerprints, which takes fully into account the problems of multiple cameras and reposting behaviors.

cs.CV