SearcharxivSearch

arXiv subjects

Xiaoqing Li

Publications and source records attributed to Xiaoqing Li.

At least 19 recordsLinked to original sources

Explainable Machine Learning for Phishing Detection on Heterogeneous Datasets with MCP-Enabled Deployment

With the growth in digital transformation and Internet usage, the Social Engineering techniques such as Phishing have become a major concern for the users and the organizations. Phishing attacks involve deceptive techniques to trick users into revealing confidential information that causes financial loss and reputation damage to organizations. According to report of Verizon, 36% of all data breaches involved phishing, highlighting the need for intelligent, adaptive, and explainable security mechanisms. This paper examines the efficiency of different machine learning algorithms in phishing detection on heterogeneous phishing datasets that include a publicly available UCI dataset, our generated datasets using tools such as EvilGinx and Zphisher, and AI generated datasets. Moreover, this work incorporates explainable AI (XAI) techniques such as Information Gain, SHAP (SHapley Additive Explanations), and LIME (Local Interpretable Model-Agnostic Explanations) to examine the most influential features impacting classification outcomes. To support practical deployment, this work also incorporates an MCP-based phishing URL detection system that offers real-time URL analysis, feature extraction, confidence-based classification, and AI-assisted security interpretation. The experimental results demonstrate that among classical models the highest accuracy is obtained by Logistic Regression at 92.44%, among ensemble models CatBoost achieved the highest accuracy at 95.01%, among neural network CNN achieved an accuracy of 94.02%, and among transformer-based models, DistilBERT got the highest accuracy at 99.78%

cs.CR

Cloud-top infrared observations reveal the four-dimensional precipitation structure

Accurate four-dimensional (4D) precipitation information is essential for understanding the Earth's energy and water cycles, yet remains observationally unresolved at global scales. Conventional theory holds that geostationary infrared observations primarily sense cloud-top properties, with limited sensitivity to sub-cloud precipitation. Here we show that cloud-top infrared measurements nevertheless encode sufficient information to recover the four-dimensional structure of precipitation, revealing a previously unexploited observability of sub-cloud processes. We introduce a physically constrained deep learning framework, 4DPrecipNet, in which a moisture-first constraint requires the latent representation to recover precipitable water vapour, anchoring the model in thermodynamic consistency. By integrating multi-channel infrared radiances with these constraints and radar-derived precipitation profiles, we reconstruct the vertical and temporal evolution of precipitation systems from geostationary orbit. The framework captures deep convective structures and their evolution, with robust performance across large samples and independent radar comparisons. These results demonstrate that sub-cloud precipitation is physically encoded in cloud-top infrared observations, establishing a new pathway for continuous global monitoring of precipitation structure.

cs.CV

Accelerated Design of Mechanically Hard Magnetically Soft High-entropy Alloys via Multi-objective Bayesian Optimization

Designing high-entropy alloys (HEAs) that are both mechanically hard and possess soft magnetic properties is inherently challenging, as a trade-off is needed for mechanical and magnetic properties. In this study, we optimize HEA compositions using a multi-objective Bayesian optimization (MOBO) framework to achieve simultaneous optimal mechanical and magnetic properties. An ensemble surrogate model is constructed to enhance the accuracy of machine learning surrogate models, while an efficient sampling strategy combining Monte Carlo sampling and acquisition function is applied to explore the high-dimensional compositional space. The implemented MOBO strategy successfully identifies Pareto-optimal compositions with enhanced mechanical and magnetic properties. The ensemble model provides robust and reliable predictions, and the sampling approach reduces the likelihood of entrapment in local optima. Our findings highlight specific elemental combinations that meet the dual design objectives, offering guidance for the synthesis of next-generation HEAs.

cond-mat.mtrl-sci

HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization

Transformers have become the de facto architecture for a wide range of machine learning tasks, particularly in large language models (LLMs). Despite their remarkable performance, many challenges remain in training deep transformer networks, especially regarding the position of the layer normalization. While Pre-Norm structures facilitate more stable training owing to their stronger identity path, they often lead to suboptimal performance compared to Post-Norm. In this paper, we propose $\textbf{HybridNorm}$, a simple yet effective hybrid normalization strategy that integrates the advantages of both Pre-Norm and Post-Norm. Specifically, HybridNorm employs QKV normalization within the attention mechanism and Post-Norm in the feed-forward network (FFN) of each transformer block. We provide both theoretical insights and empirical evidence to demonstrate that HybridNorm improves the gradient flow and the model robustness. Extensive experiments on large-scale transformer models, including both dense and sparse variants, show that HybridNorm consistently outperforms both Pre-Norm and Post-Norm approaches across multiple benchmarks. These findings highlight the potential of HybridNorm as a more stable and effective technique for improving the training and performance of deep transformer models. Code is available at https://github.com/BryceZhuo/HybridNorm.

cs.CL

A Clinical-grade Universal Foundation Model for Intraoperative Pathology

Intraoperative pathology is pivotal to precision surgery, yet its clinical impact is constrained by diagnostic complexity and the limited availability of high-quality frozen-section data. While computational pathology has made significant strides, the lack of large-scale, prospective validation has impeded its routine adoption in surgical workflows. Here, we introduce CRISP, a clinical-grade foundation model developed on over 100,000 frozen sections from eight medical centers, specifically designed to provide Clinical-grade Robust Intraoperative Support for Pathology (CRISP). CRISP was comprehensively evaluated on more than 15,000 intraoperative slides across nearly 100 retrospective diagnostic tasks, including benign-malignant discrimination, key intraoperative decision-making, and pan-cancer detection, etc. The model demonstrated robust generalization across diverse institutions, tumor types, and anatomical sites-including previously unseen sites and rare cancers. In a prospective cohort of over 2,000 patients, CRISP sustained high diagnostic accuracy under real-world conditions, directly informing surgical decisions in 92.6% of cases. Human-AI collaboration further reduced diagnostic workload by 35%, avoided 105 ancillary tests and enhanced detection of micrometastases with 87.5% accuracy. Together, these findings position CRISP as a clinical-grade paradigm for AI-driven intraoperative pathology, bridging computational advances with surgical precision and accelerating the translation of artificial intelligence into routine clinical practice.

cs.LG

Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models

Transformers have found extensive applications across various domains due to the powerful fitting capabilities. This success can be partially attributed to their inherent nonlinearity. Thus, in addition to the ReLU function employed in the original transformer architecture, researchers have explored alternative modules such as GeLU and SwishGLU to enhance nonlinearity and thereby augment representational capacity. In this paper, we propose a novel category of polynomial composition activations (PolyCom), designed to optimize the dynamics of transformers. Theoretically, we provide a comprehensive mathematical analysis of PolyCom, highlighting its enhanced expressivity and efficacy relative to other activation functions. Notably, we demonstrate that networks incorporating PolyCom achieve the $\textbf{optimal approximation rate}$, indicating that PolyCom networks require minimal parameters to approximate general smooth functions in Sobolev spaces. We conduct empirical experiments on the pre-training configurations of large language models (LLMs), including both dense and sparse architectures. By substituting conventional activation functions with PolyCom, we enable LLMs to capture higher-order interactions within the data, thus improving performance metrics in terms of accuracy and convergence rates. Extensive experimental results demonstrate the effectiveness of our method, showing substantial improvements over other activation functions. Code is available at https://github.com/BryceZhuo/PolyCom.

cs.CL

Reproducibility Assessment of Magnetic Resonance Spectroscopy of Pregenual Anterior Cingulate Cortex across Sessions and Vendors via the Cloud Computing Platform CloudBrain-MRS

Given the need to elucidate the mechanisms underlying illnesses and their treatment, as well as the lack of harmonization of acquisition and post-processing protocols among different magnetic resonance system vendors, this work is to determine if metabolite concentrations obtained from different sessions, machine models and even different vendors of 3 T scanners can be highly reproducible and be pooled for diagnostic analysis, which is very valuable for the research of rare diseases. Participants underwent magnetic resonance imaging (MRI) scanning once on two separate days within one week (one session per day, each session including two proton magnetic resonance spectroscopy (1H-MRS) scans with no more than a 5-minute interval between scans (no off-bed activity)) on each machine. were analyzed for reliability of within- and between- sessions using the coefficient of variation (CV) and intraclass correlation coefficient (ICC), and for reproducibility of across the machines using correlation coefficient. As for within- and between- session, all CV values for a group of all the first or second scans of a session, or for a session were almost below 20%, and most of the ICCs for metabolites range from moderate (0.4-0.59) to excellent (0.75-1), indicating high data reliability. When it comes to the reproducibility across the three scanners, all Pearson correlation coefficients across the three machines approached 1 with most around 0.9, and majority demonstrated statistical significance (P<0.01). Additionally, the intra-vendor reproducibility was greater than the inter-vendor ones.

stat.ML

FwNet-ECA: A Classification Model Enhancing Window Attention with Global Receptive Fields via Fourier Filtering Operations

Windowed attention mechanisms were introduced to mitigate the issue of excessive computation inherent in global attention mechanisms. In this paper, we present FwNet-ECA, a novel method that utilizes Fourier transforms paired with learnable weight matrices to enhance the spectral features of images. This method establishes a global receptive field through Filter Enhancement and avoids the use of moving window attention. Additionally, we incorporate the Efficient Channel Attention (ECA) module to improve communication between different channels. Instead of relying on physically shifted windows, our approach leverages frequency domain enhancement to implicitly bridge information across spatial regions. We validate our model on the iCartoonFace dataset and conduct downstream tasks on ImageNet, demonstrating that our model achieves lower parameter counts and computational overheads compared to shifted window approaches, while maintaining competitive accuracy. Furthermore, our visualization operations clearly demonstrated that the Filter Enhancement technique achieves greater effectiveness in the model's shallow layers, where feature maps are relatively larger. This work offers a more efficient and effective alternative for leveraging attention mechanisms in visual processing tasks, alleviating the challenges associated with windowed attention models. Code is available at https://github.com/qingxiaoli/FwNet-ECA

cs.CV

Scale-Distribution Decoupling: Enabling Stable and Effective Training of Large Language Models

Training stability is a persistent challenge in the pre-training of large language models (LLMs), particularly for architectures such as Post-Norm Transformers, which are prone to gradient explosion and dissipation. In this paper, we propose Scale-Distribution Decoupling (SDD), a novel approach that stabilizes training by explicitly decoupling the scale and distribution of the weight matrix in fully-connected layers. SDD applies a normalization mechanism to regulate activations and a learnable scaling vector to maintain well-conditioned gradients, effectively preventing $\textbf{gradient explosion and dissipation}$. This separation improves optimization efficiency, particularly in deep networks, by ensuring stable gradient propagation. Experimental results demonstrate that our method stabilizes training across various LLM architectures and outperforms existing techniques in different normalization configurations. Furthermore, the proposed method is lightweight and compatible with existing frameworks, making it a practical solution for stabilizing LLM training. Code is available at https://github.com/kaihemo/SDD.

cs.CL

High-throughput design of all-d-metal Heusler alloys for magnetocaloric applications

Due to their versatile composition and customizable properties, A$_2$BC Heusler alloys have found applications in magnetic refrigeration, magnetic shape memory effects, permanent magnets, and spintronic devices. The discovery of all-$d$-metal Heusler alloys with improved mechanical properties compared to those containing main group elements, presents an opportunity to engineer Heuslers alloys for energy-related applications. Using high-throughput density functional theory calculations, we screened magnetic all-$d$-metal Heusler compounds and identified 686 (meta)stable compounds. Our detailed analysis revealed that the inverse Heusler structure is preferred when the electronegativity difference between the A and B/C atoms is small, contrary to conventional Heusler alloys. Additionally, our calculations of Pugh ratios and Cauchy pressures demonstrated that ductile and metallic bonding are widespread in all-$d$-metal Heuslers, supporting their enhanced mechanical behaviour. We identified 49 compounds with a double-well energy surface based on Bain path calculations and magnetic ground states, indicating their potential as candidates for magnetocaloric and shape memory applications. Furthermore, by calculating the free energies, we propose that 11 compounds exhibit structural phase transitions, and propose isostructural substitution to enhance the magnetocaloric effect.

cond-mat.mtrl-sci

SABRE Enhancement with Oscillating Pulse Sequences: Symmetry Reduces Robustness

SABRE (Signal Amplification by Reversible Exchange) methods provide a simple, fast, and cost-effective method to hyperpolarize a wide variety of molecules in solution, and have been demonstrated with protons and, more recently, with heteronuclei (X-SABRE). The conventional analysis of the SABRE effect is based on level anti-crossings (LACs), which requires very low magnetic fields (~ 0.6uT) to achieve resonance and transfer spin order from the para-hydrogen to target heteronuclei. We have demonstrated in our recent study that the validity of LACs used in SABRE is very limited, so the maximum SABRE polarization predicted with LACs is not correct. Here, we present several oscillating pulse sequences that use magnetic fields far away from the resonance condition and can commonly triple the polarization. An analysis with average Hamiltonian theory indicates that the oscillating pulse, in effect, adjusts the J-couplings between hydrides and target nuclei and that a much weaker coupling produces maximum polarization. This theoretical treatment, combined with simulations and experiment, show substantial magnetization improvements relative to traditional X-SABRE methods. It also shows that, in contrast to most pulse sequence applications, waveforms with reduced time symmetry in the toggling frame make magnetization generation more robust to experimental imperfections.

physics.chem-ph

Improving SABRE hyperpolarization with highly non-intuitive pulse sequences: moving beyond avoided crossings to describe dynamics

Signal Amplification By Reversible Exchange (SABRE) creates hyperpolarization (large spin magnetization) using a transition metal catalyst and parahydrogen, addressing the sensitivity limitations of magnetic resonance. SABRE and its heteronuclear variant X-SABRE are simple, fast and general, but to date have not produced polarization levels as large as more established methods. We show here that inaccuracies in the theoretical framework for these applications, which focuses on avoided crossings (also called level anti-crossings or LACs) steer current SABRE and X-SABRE experiments away from optimal solutions. Accurate simulations show astonishingly rich and unexpected dynamics in SABRE/X-SABRE, which we explain with a combination of perturbation theory and average Hamiltonian approaches. This theoretical picture predicts simple pulse sequences with field values far from LACs (both instantaneously and on average), using different terms in the effective Hamiltonian to strategically control evolution and improve polarization transfer. Significant signal enhancements under such highly non-intuitive conditions are verified experimentally.

physics.chem-ph

Ductile and brittle crack-tip response in equimolar refractory high-entropy alloys

Understanding the strengthening and deformation mechanisms in refractory high-entropy alloys (HEAs), proposed as new high-temperature material, is required for improving their typically insufficient room-temperature ductility. Here, density-functional theory simulations and a continuum mechanics analysis were conducted to systematically investigate the competition between cleavage decohesion and dislocation emission from a crack tip in the body-centered cubic refractory HEAs HfNbTiZr, MoNbTaVW, MoNbTaW, MoNbTiV, and NbTiVZr. This crack-tip competition is evaluated for tensile loading and a totality of 15 crack configurations and slip systems. Our results predict that dislocation plasticity at the crack tip is generally unfavorable -- although the competition is close for some crack orientations, suggesting intrinsic brittleness and low crack-tip fracture toughness in these five HEAs at zero temperature. Fluctuations in local alloy composition, investigated for HfNbTiZr, can locally reduce the resistance to dislocation emission for a slip system relative to the configuration average of that slip system, but do not change the dominant crack-tip response. In the case of single-crystal MoNbTaW, where an experimental, room-temperature fracture-toughness value is available for a crack on a \{100\} plane, theoretical and experimental results agree favorably. Factors that may limit the agreement are discussed. We survey the effect of material anisotropy on preferred crack tip orientations, which are found to be alloy specific. Mixed-mode loadings are found to shift the competition in favor of cleavage or dislocation nucleation, depending on crack configuration and amplified by the effect of material anisotropy on crack tip stresses.

cond-mat.mtrl-sci

A Spectral reciprocity formula and non-vanishing for L-functions on GL(4)xGL(2)

We develop a reciprocity formula for a spectral sum over central values of L-functions on GL(4)xGL(2). As an application we show that for any self-dual cusp form Pi for SL(4,Z), there exists a Maass form pi for SL(2,Z) such that L(1/2, Pi x pi) is nonvanishing. An important ingredient is a "balanced" Voronoi summation formula involving Kloosterman sums on both sides, which can also be thought of as the functional equation of a certain double Dirichlet series involving Kloosterman sums and GL(4) Hecke eigenvalues.

math.NT

Femtosecond Photonic Viral Inactivation Probed Using Solid-State Nanopores

We report on the detection of inactivation of virus particles using femtosecond laser radiation by measuring the conductance of a solid state nanopore designed for detecting single virus particles. Conventional methods of assaying for viral inactivation based on plaque forming assays require 24-48 hours for bacterial growth. Nanopore conductance measurements provide information on morphological changes at a single virion level. We show that analysis of a time series of nanopore conductance can quantify the detection of inactivation, requiring only a few minutes from collection to analysis. Morphological changes were verified by Dynamic Light Scattering (DLS). Statistical analysis maximizing the information entropy provides a measure of the Log-reduction value. Taken together, our work provides a rapid method for assaying viral inactivation with femtosecond lasers using solid-state nanopores.

physics.bio-ph

Understanding the mechanical properties of reduced activation steels

Reduced activation ferritic/martensitic (RAFM) steels are structural materials with potential application in Generation-IV fission and fusion reactors. We use density-functional theory to scrutinize the micro-mechanical properties of the main alloy phases of three RAFM steels based on the body-centered cubic FeCrWVMn solid solution. We assess the lattice parameters and elastic properties of ferromagnetic $α$-Fe and Fe$_{91}$Cr$_{9}$, which are the main building blocks of the RAFM steels, and present a detailed analysis of the calculated alloying effects of V, Cr, Mn, and W on the mechanical properties of Fe$_{91}$Cr$_{9}$. The composition dependence of the elastic parameters is decomposed into electronic and volumetric contributions and studied for alloying levels that cover the typical intervals in RAFM steels. A linear superposition of the individual solute effects on the properties of Fe$_{91}$Cr$_{9}$ is shown to provide an excellent approximation for the \emph{ab initio} values obtained for the RAFM steels. The intrinsic ductility is evaluated through Rice's phenomenological theory using the surface and unstable stacking fault energies, and the predictions are contrasted with those obtained by empirical criteria. Alloying with V or W is found to enhance the ductility, whereas additional Cr or Mn turns the RAFM base alloys more brittle.

cond-mat.mtrl-sci

First-principles prediction of the stacking fault energy of gold at finite temperature

The intrinsic stacking fault energy (ISFE) $γ$ is a material parameter fundamental to the discussion of plastic deformation mechanisms in metals. Here, we scrutinize the temperature dependence of the ISFE of Au through accurate first-principles derived Helmholtz free energies employing both the super cell approach and the axial Ising model (AIM). A significant decrease of the ISFE with temperature, $-(36$-$39)$\,\% from 0 to 890\,K depending on the treatment of thermal expansion, is revealed, which matches the estimate based on the experimental temperature coefficient $d γ/ d T $ closely. We make evident that this decrease predominantly originates from the excess vibrational entropy at the stacking fault layer, although the contribution arising from the static lattice expansion compensates it by approximately 60\,\%. Electronic excitations are found to be of minor importance for the ISFE change with temperature. We show that the Debye model in combination with the AIM captures the correct sign but significantly underestimates the magnitude of the vibrational contribution to $γ(T)$. The hexagonal close-packed (hcp) and double hcp structures are established as metastable phases of Au. Our results demonstrate that quantitative agreement with experiments can be obtained if all relevant temperature-induced excitations are considered in first-principles modeling and that the temperature dependence of the ISFE is substantial enough to be taken into account in crystal plasticity modeling.

cond-mat.mtrl-sci

One Sentence One Model for Neural Machine Translation

Neural machine translation (NMT) becomes a new state-of-the-art and achieves promising translation results using a simple encoder-decoder neural network. This neural network is trained once on the parallel corpus and the fixed network is used to translate all the test sentences. We argue that the general fixed network cannot best fit the specific test sentences. In this paper, we propose the dynamic NMT which learns a general network as usual, and then fine-tunes the network for each test sentence. The fine-tune work is done on a small set of the bilingual training data that is obtained through similarity search according to the test sentence. Extensive experiments demonstrate that this method can significantly improve the translation performance, especially when highly similar sentences are available.

cs.CL