SearcharxivSearch

arXiv subjects

Xue Jia

Publications and source records attributed to Xue Jia.

12 recordsLinked to original sources

Clustering based on Stochastic Dominance with application for risk averters and risk seekers

Stochastic Dominance (SD) theory provides a rigorous framework for selecting superior assets tailored to the asset allocation needs of investors with varying risk preferences (i.e., risk-averse, risk-seeking, and risk-neutral). However, traditional stock clustering methods typically rely on geometric metrics such as Euclidean distance, which often fail to effectively capture the intrinsic risk dominance relationships among assets. To address this limitation, this paper proposes an innovative clustering analysis framework based on SD test statistics. Methodologically, this study deeply integrates SD theory with machine learning algorithms. Transcending the limitations of traditional reliance on geometric distance, we innovatively utilize test statistics from first-, second-, and third-order SD to construct a "Stochastic Dominance Coefficient Matrix." Building upon this matrix, we modify the classic K-means and Hierarchical Clustering algorithms. Specifically, we derive 12 distinct algorithm variants tailored to different orders of SD relationships. Simultaneously, we construct the SD-SC coefficient and the SD-DBI index as specialized validity indices to evaluate the clustering performance. Empirically, we analyze constituent stock data from a representative developed market (the US NASDAQ Index) and an emerging market (China's CSI 100 Index). The results verify the effectiveness and robustness of the proposed method. Furthermore, we apply the clustering results to the modification of the Single Index Model and the construction of Global Minimum Variance Portfolios (GMVP). The findings demonstrate that the proposed method effectively facilitates customized asset allocation for investors, holding significant theoretical value and practical implications.

stat.ML

Building a physics-aware AI ecosystem for solid-state hydrogen storage materials

Hydrogen storage remains a central bottleneck for scalable hydrogen energy systems due to the multiscale and coupled nature of the thermodynamics, kinetics, and microstructural evolution of hydrogen storage materials (HSMs). Although artificial intelligence (AI) has accelerated materials discovery, current approaches remain constrained by fragmented data, limited physical consistency, and weak integration with experimental validation. Here, we propose a unified framework that integrates coherent data infrastructure, physics-grounded modeling, and AI-driven inverse design within a closed-loop discovery paradigm. By embedding physical constraints and experimental feedback, this approach enables adaptive, physically consistent optimization, thereby establishing a pathway toward autonomous, digital-twin-enabled discovery of HSMs.

cond-mat.mtrl-sci

How is a gas sensor poisoned by volatile methylsiloxanes?

Volatile methyl siloxanes (VMSs), widely present in consumer and industrial products, have attracted increasing concerns due to their persistence, bioaccumulation behavior, and adverse health effects. Beyond their environmental implications, VMSs also pose operational challenges for sensing technologies because they readily decompose on sensing materials to form silicon-based compounds (e.g., silica and silane) that irreversibly impair sensing performance, a phenomenon commonly known as siloxane poisoning. Despite its prevalence, the mechanistic basis of this deactivation remains poorly understood. Herein, we present the first comprehensive theoretical study of siloxane-induced poisoning in catalytic gas sensors. Guided by our self-developed AI Agent, Digital Sensor Platform (DigSen), we first identify siloxane poisoning as a previously overlooked yet high-impact research direction. Using hexamethyldisiloxane (HMDS) as a model compound, we then conducted first-principles calculations to uncover decomposition pathways across noble metal surfaces. Strikingly, a descriptor-based microkinetic volcano model is developed to capture the trade-off between sensing activity and resistance to poisoning, enabling predictive identification of anti-poisoning candidates. These insights not only elucidate the origin of siloxane poisoning but also demonstrate how AI-driven discovery, mechanistic theory, and experiments can be integrated into a closed-loop framework for catalytic sensor design. More broadly, this AI-guided paradigm represents a generalizable strategy for materials digital discovery, offering a transferable methodology that extends well beyond siloxane systems to diverse classes of materials challenges.

physics.chem-ph

A unified descriptor framework for hydrogen storage capacity and equilibrium pressure in interstitial hydrides

Hydrogen is a promising energy carrier, yet its practical deployment is limited by the lack of storage materials that simultaneously achieve high storage capacity ($w$) and practical equilibrium pressure at room temperature ($P_{\rm eq,RT}$). Interstitial metal hydrides offer fast kinetics and favorable thermodynamics (high $P_{\rm eq,RT}$) but suffer from intrinsically low w. Here, we establish a physically interpretable, data-driven framework to uncover descriptor-property relationships in interstitial hydrides using a curated database of pressure-composition-temperature measurements (Digital Hydrogen Platform, DigHyd) and white-box symbolic regression. Strikingly, the analysis reveals a clear separation of governing mechanisms, in which $w$ is governed by geometric and lattice conditions, captured by the average atomic radius ($\left\langle r_M \right\rangle$) and average thermal conductivity ($\left\langle\kappa\right\rangle$), with an optimal regime of $r_M \sim 1.47 \r{A}$ and relatively low $\left\langle\kappa\right\rangle$. In contrast, $P_{\rm eq,RT}$ is governed by elastic properties, captured by the average shear modulus ($\left\langle G \right\rangle$) and average Poisson's ratio ($\left\langle \nu \right\rangle$), reflecting the role of lattice rigidity and mechanical compliance. These relationships are translated into compositional optimization pathways that follow the descriptor trends above, enabling the design of candidate materials with enhanced w under practical equilibrium conditions ($P_{\rm eq,RT} \sim 0.1$ MPa). This work establishes a general, interpretable strategy for physics-informed design of energy materials systems.

cond-mat.mtrl-sci

Digital Hydrogen Platform (DigHyd): A Rigorously Curated Database for Hydrogen Storage Materials Empowered by AI-Assisted Literature Mining

Solid-state hydrogen storage materials are promising candidates for safe and compact hydrogen storage; however, data-driven discovery in this field remains limited by the availability of large-scale, well-curated datasets. Here, we present the Digital Hydrogen Platform (DigHyd: www.dighyd.org), a rigorously curated database comprising $>4,000$ experimental literature sources and $>30,000$ data entries on hydrogen storage materials, constructed through AI-assisted literature mining combined with human-in-the-loop validation. In addition to gravimetric hydrogen storage density ($w$), DigHyd also covers thermodynamic parameters, specifically the enthalpy ($\Delta H$) and entropy ($\Delta S$) changes associated with hydrogenation reactions, primarily defined as $M + \frac{1}{2} {\rm H}_2 \rightleftarrows M{\rm H}$. These parameters were obtained by manually analyzing multi-temperature pressure-composition-temperature (PCT) data using van't Hoff analysis. By focusing on $\Delta H$ and $\Delta S$ rather than fixing equilibrium pressure at a single temperature, DigHyd enables flexible evaluation of equilibrium behavior under application-specific operating conditions. Statistical analyses reveal distinct distributions of thermodynamic parameters across material classes, together with broad compositional variability within representative hydride systems. Furthermore, both physically interpretable symbolic regression and black-box XGBoost models achieve comparable predictive performance for $w$ and equilibrium pressure at room temperature ($P_{\rm eq,RT}$), demonstrating internal consistency and learnable composition-property relationships within the curated dataset. Overall, DigHyd provides a rigorously curated thermodynamic dataset that serves as a reliable basis for data-driven analyses of hydrogen storage materials and supports systematic exploration of structure-property relationships.

cond-mat.mtrl-sci

Heteronuclear and Homonuclear Vector Solitons in Lasers

Vector solitons (VSs), being observed across various fields from optics to Bose-Einstein condensates, are localized structures composed of orthogonal modes bound by nonlinear couplings. Nevertheless, the influence of intermodal linear coupling on the physical properties of this bimodal structure remains to be decently revealed and harnessed. Utilizing an ultrafast fiber laser as a platform, we predict and demonstrate that the linear mode coupling (LMC) induces the deformable VS in terms of the temporal and spectral structures. Weak LMC supports heteronuclear vector solitons built of dissimilar polarization modes, i.e., a single pulse coupled to an orthogonal damped pulse chain. On the other hand, strong LMC facilitates the homonuclear VS composed of polarization modes with similar structures, in the form of soliton compounds featuring caterpillar motions. Our findings reveal new patterns of VSs and open an effective avenue for versatile ultrafast optical sources.

physics.optics

Physically Interpretable Descriptors Drive the Materials Design of Metal Hydrides for Hydrogen Storage

Designing metal hydrides for hydrogen storage remains a longstanding challenge due to the vast compositional space and complex structure-property relationships. Herein, for the first time, we present physically interpretable models for predicting two key performance metrics, gravimetric hydrogen density $w$ and equilibrium pressure $P_{\rm eq,RT}$ at room temperature, based on a minimal set of chemically meaningful descriptors. Using a rigorously curated dataset of $5,089$ metal hydride compositions from our recently developed Digital Hydrogen Platform (\it{DigHyd}) based on large-scale data mining from available experimental literature of solid-state hydrogen storage materials, we systematically constructed over $1.6$ million candidate models using combinations of scalar transformations and nonlinear link functions. The final closed-form models, derived from $2$-$3$ descriptors each, achieve predictive accuracies on par with state-of-the-art machine learning methods, while maintaining full physical transparency. Strikingly, descriptor-based design maps generated from these models reveal a fundamental trade-off between $w$ and $P_{\rm eq,RT}$: saline-type hydrides, composed of light electropositive elements, offer high $w$ but low $P_{\rm eq,RT}$, whereas interstitial-type hydrides based on heavier electronegative transition metals show the opposite trend. Notably, Be-based systems, such as Be-Na alloys, emerge as rare candidates that simultaneously satisfy both performance metrics, attributed to the unique combination of light mass and high molar density for Be. Our models indicate that Be-based systems may offer renewed prospects for approaching these benchmarks. These results provide chemically intuitive guidelines for materials design and establish a scalable framework for the rational discovery of materials in complex chemical spaces.

cond-mat.mtrl-sci

"DIVE" into Hydrogen Storage Materials Discovery with AI Agents

Data-driven artificial intelligence (AI) approaches are fundamentally transforming the discovery of new materials. Despite the unprecedented availability of materials data in the scientific literature, much of this information remains trapped in unstructured figures and tables, hindering the construction of large language model (LLM)-based AI agent for automated materials design. Here, we present the Descriptive Interpretation of Visual Expression (DIVE) multi-agent workflow, which systematically reads and organizes experimental data from graphical elements in scientific literatures. We focus on solid-state hydrogen storage materials-a class of materials central to future clean-energy technologies and demonstrate that DIVE markedly improves the accuracy and coverage of data extraction compared to the direct extraction by multimodal models, with gains of 10-15% over commercial models and over 30% relative to open-source models. Building on a curated database of over 30,000 entries from 4,000 publications, we establish a rapid inverse design workflow capable of identifying previously unreported hydrogen storage compositions in two minutes. The proposed AI workflow and agent design are broadly transferable across diverse materials, providing a paradigm for AI-driven materials discovery.

cs.AI

A Materials Map Integrating Experimental and Computational Data via Graph-Based Machine Learning for Enhanced Materials Discovery

Materials informatics (MI), emerging from the integration of materials science and data science, is expected to significantly accelerate material development and discovery. The data used in MI are derived from both computational and experimental studies; however, their integration remains challenging. In our previous study, we reported the integration of these datasets by applying a machine learning model that is trained on the experimental dataset to the compositional data stored in the computational database. In this study, we use the obtained datasets to construct materials maps, which visualize the relationships between material properties and structural features, aiming to support experimental researchers. The materials map is constructed using the MatDeepLearn (MDL) framework, which implements materials property prediction using graph-based representations of material structure and deep learning modeling. Through statistical analysis, we find that the MDL framework using the message passing neural network (MPNN) architecture efficiently extracts features reflecting the structural complexity of materials. Moreover, we find that this advantage does not necessarily translate into improved accuracy in the prediction of material properties. We attribute this unexpected outcome to the high learning performance inherent in MPNN, which can contribute to the structuring of data points within the materials map.

cond-mat.mtrl-sci

Universal Catalyst Design Framework for Electrochemical Hydrogen Peroxide Synthesis Facilitated by Local Atomic Environment Descriptors

Developing a universal and precise design framework is crucial to search high-performance catalysts, but it remains a giant challenge due to the diverse structures and sites across various types of catalysts. To address this challenge, herein, we developed a novel framework by the refined local atomic environment descriptors (i.e., weighted Atomic Center Symmetry Function, wACSF) combined with machine learning (ML), microkinetic modeling, and computational high-throughput screening. This framework is successfully integrated into the Digital Catalysis Database (DigCat), enabling efficient screening for 2e- water oxidation reaction (2e- WOR) catalysts across four material categories (i.e., metal alloys, metal oxides and perovskites, and single-atom catalysts) within a ML model. The proposed wACSF descriptors integrating both geometric and chemical features are proven effective in predicting the adsorption free energies with ML. Excitingly, based on the wACSF descriptors, the ML models accurately predict the adsorption free energies of hydroxyl (ΔGOH*) and oxygen (ΔGO*) for such a wide range of catalysts, achieving R2 values of 0.84 and 0.91, respectively. Through density functional theory calculations and microkinetic modeling, a universal 2e- WOR microkinetic volcano model was derived with excellent agreement with experimental observations reported to date, which was further used to rapidly screen high-performance catalysts with the input of ML-predicted ΔGOH*. Most importantly, this universal framework can significantly improve the efficiency of catalyst design by considering multiple types of materials at the same time, which can dramatically accelerate the screening of high-performance catalysts.

cond-mat.mtrl-sci

Generator polynomials of cyclic expurgated or extended Goppa codes

Classical Goppa codes are a well-known class of codes with applications in code-based cryptography, which are a special case of alternant codes. Many papers are devoted to the search for Goppa codes with a cyclic extension or with a cyclic parity-check subcode. Let $\Bbb F_q$ be a finite field with $q=2^l$ elements, where $l$ is a positive integer. In this paper, we determine all the generator polynomials of cyclic expurgated or extended Goppa codes under some prescribed permutations induced by the projective general linear automorphism $A \in PGL_2(\Bbb F_q)$. Moreover, we provide some examples to support our findings.

cs.IT

Determining hulls of generalized Reed-Solomon codes from algebraic geometry codes

In this paper, we provide conditions that hulls of generalized Reed-Solomon (GRS) codes are also GRS codes from algebraic geometry codes. If the conditions are not satisfied, we provide a method of linear algebra to find the bases of hulls of GRS codes and give formulas to compute their dimensions. Besides, we explain that the conditions are too good to be improved by some examples. Moreover, we show self-orthogonal and self-dual GRS codes.

cs.IT