SearcharxivSearch

arXiv subjects

Xiaoqing Liu

Publications and source records attributed to Xiaoqing Liu.

At least 19 recordsLinked to original sources

OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization

NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still degrade quantization accuracy within NVFP4 blocks. Within each quantization block, large activations can dominate the block scale, increasing the quantization error of the remaining values sharing the same scale. Existing post-training quantization (PTQ) methods mitigate outlier errors through strategies such as mixed precision, rotation, or residual compensation, but these approaches are either not specifically tailored to NVFP4 or introduce additional computation. In this work, we revisit NVFP4 from a channel-grouping perspective and define the reducible error incurred by remaining block values under the scale set by the block maximum as Collateral Quantization Error. Based on this insight, we propose OCGQuant, a post-training quantization method centered on Outlier-Companion Grouping (OCG), which adaptively pairs outlier channels with low-magnitude companion channels to improve NVFP4 activation block composition. Experiments on Llama3 and Qwen3 show that OCGQuant achieves the lowest WikiText-2 perplexity and highest average downstream accuracy among evaluated PTQ methods, while maintaining prefill speedup close to RTN and matching its peak decoding memory. Code is available at https://github.com/Eshamont/OCGQuant.

cs.CL

Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion

Multimodal learning faces modality imbalance, where dominant modalities suppress weaker ones due to inconsistent convergence rates. Existing static or heuristic methods overlook sample-level variations in prediction bias and fail to isolate low-quality outlier samples. To address this, we propose a novel framework to quantitatively diagnose and dynamically mitigate modality imbalance at the sample level. We first introduce a Modality Gap metric to quantify prediction discrepancies between unimodal branches. Empirical analysis reveals a distinct bimodal distribution, reflecting the natural coexistence of balanced and imbalanced sample subgroups. We then employ a Gaussian Mixture Model (GMM) to model this gap distribution, leveraging Bayesian posterior probabilities for probabilistic soft separation of subgroups. Next, we construct a two-stage training framework comprising a Warm-up stage and an Adaptive Training stage. In the Adaptive Training stage, a GMM-guided Adaptive Loss dynamically reallocates optimization priorities, imposing stronger modality alignment penalties on imbalanced samples while prioritizing multimodal fusion for balanced ones. Experimental results demonstrate that our method significantly outperforms current state-of-the-art baselines. Furthermore, fine-tuning on a high-quality balanced subset filtered by the GMM serves as an effective data purification strategy.

cs.LG

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

High-quality data is scarce in large language model (LLM) training, yet how to schedule its use with optimization dynamics lacks theoretical guidance. We extend functional scaling laws with time-varying data quality and derive asymptotically optimal joint data-quality and batch-size schedules within a feature-space regression model. The solution reveals two regimes and dual uses of high-quality data: in the noise-limited regime, a smaller batch converts cleaner data into more signal at comparable noise; in the signal-limited regime, late placement suppresses terminal noise without sacrificing signal accumulation. This explains why conventional decay schedules can conflict with curriculum-style pipelines. Motivated by the theoretical structure, we propose Drop-Stable-Rampup for LLM midtraining: drop the batch size at the quality transition, keep it low to accumulate signal, then ramp up to suppress noise. On a 15B MoE model midtrained on 108B tokens of general-domain proprietary data, Drop-Stable-Rampup improves average accuracy over Warmup-Stable-Decay by +1.70 and Cosine-decay by +2.98, including +4.23 on GSM8K and +2.80 on MATH. On a public math-and-code mixture, it leads all reported STEM, mathematics, and code benchmarks, improving the overall mean over the strongest baseline by +3.27 on a 600M dense model and +5.25 on the same MoE architecture.

cs.LG

Self-Synchronized Terahertz and X-Ray Free-Electron Lasers from a Single Pre-Bunched Electron Beam

Ultrafast pump-probe spectroscopy combining intense terahertz (THz) and X-ray pulses is a critical tool for investigating complex structural and electronic dynamics in materials. However, current setups combining THz sources and X-ray free-electron lasers (FELs) often suffer from high system complexity, inherent timing jitter, or limited THz pulse properties. Here, we experimentally demonstrate the generation of intrinsically synchronized, strong-field, narrow-band THz and X-ray FELs from a single pre-bunched electron beam. Sequentially passing the beam through X-ray and THz amplifiers reveals a highly synergistic process: the initial periodic THz density modulation notably boosts the X-ray FEL pulse energy, while robustly surviving the intense X-ray emission to drive high-power, narrow-band THz radiation. Originating from the same electron bunch, the two pulses inherently maintain a precise, constant time delay. This jitter-free scheme establishes a highly reliable platform tailored for both X-ray-pump/THz-probe and THz-pump/X-ray-probe experiments.

physics.acc-ph

HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bench

SWE-bench has emerged as the premier benchmark for evaluating Large Language Models on complex software engineering tasks. While these capabilities are fundamentally acquired during the mid-training phase and subsequently elicited during Supervised Fine-Tuning (SFT), there remains a critical deficit in metrics capable of guiding mid-training effectively. Standard metrics such as Perplexity (PPL) are compromised by the "Long-Context Tax" and exhibit weak correlation with downstream SWE performance. In this paper, we bridge this gap by first introducing a rigorous data filtering strategy. Crucially, we propose the Entropy Compression Hypothesis, redefining intelligence not by scalar Top-1 compression, but by the capacity to structure uncertainty into Entropy-Compressed States of low orders ("reasonable hesitation"). Grounded in this fine-grained entropy analysis, we formulate a novel metric, HE-SNR (High-Entropy Signal-to-Noise Ratio). We validate our approach on models with up to 560B parameters across different context windows (32K/128K). This work provides both the theoretical foundation and practical tools for optimizing the latent potential of LLMs in complex engineering domains.

cs.LG

Fully coherent short wavelength free-electron laser driven by a single sub-microjoule seed

High-repetition-rate, fully coherent extreme-ultraviolet (EUV) and X-ray free-electron lasers (FELs) are essential for advanced time-resolved ultrafast spectroscopies. While external seeding serves as the standard technique to achieve precise temporal coherence, conventional methods demand hundred-megawatt peak-power laser systems. Furthermore, advanced configurations like echo-enabled harmonic generation (EEHG) introduce the severe complexities of dual-laser synchronization. Together, these requirements fundamentally restrict operations to kilohertz repetition rates and compromise overall system stability. Here, we experimentally demonstrate a fully coherent EEHG-FEL driven by a single, sub-microjoule seed laser. By employing a direct-amplification enabled harmonic generation technique, we utilize an initial 0.4 microJ (2 MW peak power) ultraviolet seed to directly drive coherent lasing at nanometer wavelengths. By eliminating the need for extreme peak powers and multiple synchronized lasers, this approach significantly simplifies the seeding architecture and provides a practical and robust pathway toward megahertz-class, fully coherent EUV and X-ray light sources.

physics.acc-ph

Meta-LegNet: A Transferable and Interpretable Framework for Surface Adsorption Prediction via Self-Defined Adsorption-Environment Learning

A central challenge in computational catalysis is the identification of low-energy and chemically plausible adsorption configurations, as these directly affect adsorption energies, reaction pathways, and catalytic performance. Existing approaches generally rely on enumerating candidate adsorption sites followed by iterative refinement through density functional theory calculations or machine-learning-based relaxations. However, such workflows remain computationally expensive and are difficult to scale to complex surfaces or multi-adsorbate systems. Here, we introduce Meta-LegNet, a graph learning framework that combines SE(3)-equivariant atom-level message passing with voxel-based multiscale aggregation and cross-domain meta-learning to learn transferable representations of local adsorption environments across diverse catalyst--adsorbate systems. Rather than following a conventional regression-only paradigm, Meta-LegNet encodes local chemical environments using invariant radial features and equivariant directional information, and further incorporates broader structural context through coordinate-frame voxel pooling, assignment-based upsampling, and gated feature fusion. The resulting local-global decomposition produces atom-resolved attribution maps, which are processed to identify adsorption-relevant local environments in an interpretable manner. Based on the learned representations, we further construct an adsorption-environment database and develop a template-matching strategy to propose likely adsorption sites on previously unexplored surfaces without exhaustive site enumeration. Overall, our results suggest that learning transferable adsorption environments provides an accurate, interpretable, and practical route for accelerating catalyst screening.

cond-mat.mtrl-sci

An Efficient High-Degree, High-Order Equivariant Graph Neural Network for Direct Crystal Structure Optimization

Crystal structure optimization is fundamental to materials modeling but remains computationally expensive when performed with density-functional theory (DFT). Machine-learning (ML) approaches offer substantial acceleration, yet existing methods face three key limitations: (i) most models operate solely on atoms and treat lattice vectors implicitly, despite their central role in structural optimization; (ii) they lack efficient mechanisms to capture high-degree angular information and higher-order geometric correlations simultaneously, which are essential for distinguishing subtle structural differences; and (iii) many pipelines are multi-stage or iterative rather than truly end-to-end, making them prone to error accumulation and limiting scalability. Here we present E$^{3}$Relax-H$^{2}$, an end-to-end high-degree, high-order equivariant graph neural network that maps an initial crystal directly to its relaxed structure. The key idea is to promote both atoms and lattice vectors to graph nodes, enabling a unified and symmetry-consistent representation of structural degrees of freedom. Building on this formulation, E$^{3}$Relax-H$^{2}$ introduces two message-passing mechanisms: (i) a high-degree, high-order message-passing module that efficiently captures high-degree angular representations and high-order many-body correlations; and (ii) a lattice-atom message-passing module that explicitly models the bidirectional coupling between lattice deformation and atomic displacement. In addition, we propose a differentiable periodicity-aware Cartesian displacement loss tailored for one-shot structure prediction under periodic boundary conditions.

cond-mat.mtrl-sci

Demonstration of High-Gain Harmonic Lasing in a Terahertz Free-Electron Laser

Compact Free-Electron Lasers (FELs) offering broad, continuous spectral tunability are traditionally constrained by fixed-parameter magnetic structures and the necessity for high-energy electron beams. High-gain Harmonic Lasing (HL) has long been proposed as a solution to overcome these limitations; however, a robust experimental verification of this principle has remained absent. Here, we report the first experimental demonstration of high-gain HL. By employing a frequency-tunable electron beam density modulation to dominate the fundamental instability, we achieved sustained FEL amplification at the 3rd and 5th harmonics of the wiggler. The HL mode generated output power comparable to conventional fundamental operation with enhanced stability and narrower spectral bandwidth. Notably, we demonstrate that HL extends the spectral coverage by a factor of two under fixed facility constraints, achieving pulse energies up to 540 μJ. These results establish high-gain HL as a versatile mechanism for advancing compact, wavelength-flexible FEL facilities.

physics.acc-ph

An AI-ready fine-tuning framework for accurate machine-learning interatomic potentials in solid-solid battery interfaces

Atomistic modeling of solid-solid battery interfaces is essential for understanding electro-chemo-mechanical coupling, but the complex interfacial chemistry and heterogeneous environments pose major challenges for quantum-accurate, data-efficient modeling. Herein, we propose an approach of fine-tuning with integrated replay and efficiency (FIRE), a general framework for universal machine-learning interatomic potentials by combining efficient configurational sampling with a replay-argumented continual strategy, achieving quantum-level accuracy at moderate cost. Across six solid-solid battery interface systems, FIRE consistently achieves root-mean-square errors in energy below 1 meV/atom and in force near 20 meV/angstrom, marking an order-of-magnitude improvement over existing models while requiring only 10% of the original datasets. In addition, the fine-tuned model successfully reproduces key mechanical and electrochemical properties of the materials, in close agreement with experimental data. The FIRE offers a generalizable and data-efficient approach for developing accurate interatomic potentials across diverse materials, enabling predictive simulations beyond the reach of first-principles methods.

cond-mat.mtrl-sci

Beyond Adam: Disentangling Optimizer Effects in the Fine-Tuning of Atomistic Foundation Models

Atomistic foundation models constitute a paradigm shift in computational materials science by providing universal machine-learned interatomic potentials with broad transferability across chemical spaces. Although fine-tuning is essential for adapting these pretrained models to specific target systems, the influence of the optimization algorithm on this process remains insufficiently characterized. In this work, we perform a rigorous benchmark of seven first-order optimizers, including Adam, AdamW, RAdam, SGD, LAMB, Ranger, and ScheduleFree, for the fine-tuning of foundation models across molecular, crystalline, and liquid regimes. We evaluate these algorithms based on energy and force accuracy for both in-distribution and out-of-distribution configurations, as well as their impact on downstream physical properties such as elastic moduli, phonon spectra, and interfacial dynamics. We interpret these empirical results through a preconditioning framework that views each optimizer as a data-dependent linear transformation of the gradient. This analysis clarifies how different update rules impose specific spectral filters on the effective loss Hessian. Across all regimes, AdamW and ScheduleFree achieve superior curvature conditioning and force accuracy, whereas stochastic gradient descent exhibits slow convergence and instability. Furthermore, we demonstrate that a brief second-order refinement stage reduces residual anisotropy in the loss landscape and enhances the fidelity of physical observables without increasing inference costs. These findings provide conceptual insight and practical guidance for selecting and designing optimizers to ensure the stable and efficient fine-tuning of universal interatomic potentials.

physics.comp-ph

Equivariant Atomic and Lattice Modeling Using Geometric Deep Learning for Crystal Structure Optimization

Structure optimization, which yields the relaxed structure (minimum-energy state), is essential for reliable materials property calculations, yet traditional ab initio approaches such as density-functional theory (DFT) are computationally intensive. Machine learning (ML) has emerged to alleviate this bottleneck but suffers from two major limitations: (i) existing models operate mainly on atoms, leaving lattice vectors implicit despite their critical role in structural optimization; and (ii) they often rely on multi-stage, non-end-to-end workflows that are prone to error accumulation. Here, we present E3Relax, an end-to-end equivariant graph neural network that maps an unrelaxed crystal directly to its relaxed structure. E3Relax promotes both atoms and lattice vectors to graph nodes endowed with dual scalar-vector features, enabling unified and symmetry-preserving modeling of atomic displacements and lattice deformations. A layer-wise supervision strategy forces every network depth to make a physically meaningful refinement, mimicking the incremental convergence of DFT while preserving a fully end-to-end pipeline. We evaluate E3Relax on four benchmark datasets and demonstrate that it achieves remarkable accuracy and efficiency. Through DFT validations, we show that the structures predicted by E3Relax are energetically favorable, making them suitable as high-quality initial configurations to accelerate DFT calculations.

cond-mat.mtrl-sci

An Automatic Detection Method for Hematoma Features in Placental Abruption Ultrasound Images Based on Few-Shot Learning

Placental abruption is a severe complication during pregnancy, and its early accurate diagnosis is crucial for ensuring maternal and fetal safety. Traditional ultrasound diagnostic methods heavily rely on physician experience, leading to issues such as subjective bias and diagnostic inconsistencies. This paper proposes an improved model, EH-YOLOv11n (Enhanced Hemorrhage-YOLOv11n), based on small-sample learning, aiming to achieve automatic detection of hematoma features in placental ultrasound images. The model enhances performance through multidimensional optimization: it integrates wavelet convolution and coordinate convolution to strengthen frequency and spatial feature extraction; incorporates a cascaded group attention mechanism to suppress ultrasound artifacts and occlusion interference, thereby improving bounding box localization accuracy. Experimental results demonstrate a detection accuracy of 78%, representing a 2.5% improvement over YOLOv11n and a 13.7% increase over YOLOv8. The model exhibits significant superiority in precision-recall curves, confidence scores, and occlusion scenarios. Combining high accuracy with real-time processing, this model provides a reliable solution for computer-aided diagnosis of placental abruption, holding significant clinical application value.

cs.CV

Fine-Tuning Universal Machine-Learned Interatomic Potentials: A Tutorial on Methods and Applications

Universal machine-learned interatomic potentials (U-MLIPs) have demonstrated broad applicability across diverse atomistic systems but often require fine-tuning to achieve task-specific accuracy. While the number of available U-MLIPs and their fine-tuning applications is rapidly expanding, there remains a lack of systematic guidance on how to effectively fine-tune these models. This tutorial provides a comprehensive, step-by-step guide to fine-tuning U-MLIPs for computational materials modeling. Using the recently released MACE-MP-0 as a representative foundation model, we illustrate the full workflow of dataset preparation, hyperparameter selection, model training, and validation. Beyond methodological guidance, we conduct systematic case studies on solid-state electrolytes, stacking fault defects in metals, semiconductors, solid-liquid interfacial interactions in low-dimensional systems, and more complicated heterointerfaces. These examples demonstrate that fine-tuning substantially improves predictive accuracy while maintaining affordable computational cost, accelerates training convergence, enhances out-of-distribution generalization, and achieves superior data efficiency. Remarkably, fine-tuned foundation models can even capture aspects of long-range physics without explicit corrections. Together, these results highlight that fine-tuning not only provides a practical recipe for applying U-MLIPs, but also offers new insights into their physical fidelity and potential for advancing large-scale atomistic simulations. To support practical applications, we include code examples that enable researchers, particularly those new to the field, to efficiently incorporate fine-tuned U-MLIPs into their workflows.

physics.comp-ph

A Study on the Fine-Tuning Performance of Universal Machine-Learned Interatomic Potentials (U-MLIPs)

Universal machine-learned interatomic potentials (U-MLIPs) have demonstrated effectiveness across diverse atomistic systems but often require fine-tuning for task-specific accuracy. We investigate the fine-tuning of two MACE-based foundation models, MACE-MP-0 and its variant MACE-MP-0b, and identify key insights. Fine-tuning on task-specific datasets enhances accuracy and, in some cases, outperforms models trained from scratch. Additionally, fine-tuned models benefit from faster convergence due to the strong initial predictions provided by the foundation model. The success of fine-tuning also depends on careful dataset selection, which can be optimized through filtering or active learning. We further discuss practical strategies for achieving better fine-tuning foundation models in atomistic simulations and explore future directions for their development and applications.

physics.comp-ph

Modeling crystal defects using defect-informed neural networks

Most AI-for-Materials research to date has focused on ideal crystals, whereas real-world materials inevitably contain defects that play a critical role in modern functional technologies. The defects break geometric symmetry and increase interaction complexity, posing particular challenges for traditional ML models. Here, we introduce Defect-Informed Equivariant Graph Neural Network (DefiNet), a model specifically designed to accurately capture defect-related interactions and geometric configurations in point-defect structures. DefiNet achieves near-DFT-level structural predictions in milliseconds using a single GPU. To validate its accuracy, we perform DFT relaxations using DefiNet-predicted structures as initial configurations and measure the residual ionic steps. For most defect structures, regardless of defect complexity or system size, only 3 ionic steps are required to reach the DFT-level ground state. Finally, comparisons with scanning transmission electron microscopy (STEM) images confirm DefiNet's scalability and extrapolation beyond point defects, positioning it as a valuable tool for defect-focused materials research.

cond-mat.mtrl-sci

RoHyDR: Robust Hybrid Diffusion Recovery for Incomplete Multimodal Emotion Recognition

Multimodal emotion recognition analyzes emotions by combining data from multiple sources. However, real-world noise or sensor failures often cause missing or corrupted data, creating the Incomplete Multimodal Emotion Recognition (IMER) challenge. In this paper, we propose Robust Hybrid Diffusion Recovery (RoHyDR), a novel framework that performs missing-modality recovery at unimodal, multimodal, feature, and semantic levels. For unimodal representation recovery of missing modalities, RoHyDR exploits a diffusion-based generator to generate distribution-consistent and semantically aligned representations from Gaussian noise, using available modalities as conditioning. For multimodal fusion recovery, we introduce adversarial learning to produce a realistic fused multimodal representation and recover missing semantic content. We further propose a multi-stage optimization strategy that enhances training stability and efficiency. In contrast to previous work, the hybrid diffusion and adversarial learning-based recovery mechanism in RoHyDR allows recovery of missing information in both unimodal representation and multimodal fusion, at both feature and semantic levels, effectively mitigating performance degradation caused by suboptimal optimization. Comprehensive experiments conducted on two widely used multimodal emotion recognition benchmarks demonstrate that our proposed method outperforms state-of-the-art IMER methods, achieving robust recognition performance under various missing-modality scenarios. Our code will be made publicly available upon acceptance.

cs.CV

First Lasing and Stable Operation of a Direct-Amplification Enabled Harmonic Generation Free-Electron laser

Seeded free-electron lasers (FELs) capable of operating at repetition rates up to the MHz level are in high demand for advanced time-resolved spectroscopies, which require both full longitudinal coherence and high average photon flux in the extreme ultraviolet (EUV) and x-ray regimes. However, conventional external-seed laser systems cannot sustain MHz operation with sufficient hundreds of megawatts peak power requirement due to their limited total power. Here, we report the first lasing and stable operation of a direct-amplification-enabled harmonic generation FEL driven by a weak seed laser with MW-level peak power. Beginning with an ultraviolet seed laser with only 0.75 μJ pulse energy, we demonstrate its direct amplification to over 10 μJ within an 8-meter-long modulator. We observe coherent harmonic generation up to the 12th harmonic of the seed and achieve saturation of the 7th harmonic in the radiator. These results represent a crucial milestone toward the realization of MHz-class, fully coherent EUV and x-ray light sources.

physics.acc-ph