SearcharxivSearch

arXiv subjects

Mingjie Liao

Publications and source records attributed to Mingjie Liao.

9 recordsLinked to original sources

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

The composition of training data, governed by the diversity of sources and their mixing strategy, is a cornerstone of Large Language Model (LLM) pre-training. Online Data Mixing (ODM), the technique of adaptively adjusting data mixtures during training, has emerged as a promising direction to improve efficiency. However, existing methods are constrained by their reliance on a singular optimization perspective, which fundamentally overlooks the need for complex LLM pre-training to consider the dynamic data composition from multiple dimensions. To overcome this limitation, we introduce the Holistic Data Scheduler (HDS), a novel online data mixing framework. HDS formulates the data scheduling challenge as a reinforcement learning problem in a continuous control space and leverages the Soft Actor-Critic (SAC) algorithm for its stability and sample efficiency in exploring the high-dimensional policy space. At the core of HDS lies a novel multi-objective, holistic reward function that integrates three critical perspectives: a data-driven reward for quality, a loss-driven reward capturing inter-domain influence, and a model-driven reward based on weight norms. To validate our design and determine its optimal configuration, we conducted systematic experiments on LLMs of various sizes. On The Pile benchmark, HDS reaches the final validation perplexity of the next best method with 44% fewer training iterations. Furthermore, it achieves a 7.2% improvement on the MMLU 0-shot task along with consistent gains on other benchmarks, showcasing its ability to enhance both training efficiency and final model capability.

cs.LG

AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining

Optimizing pretraining data composition is pivotal for LLM generalization. While dynamic mixing outperforms static strategies by capturing evolving training dynamics, current methods fail to reconcile computational efficiency with sample efficiency and structural flexibility for diverse pipelines.We introduce Actor--Critic Online Data Mixing (AC-ODM), which approaches data mixing from a reinforcement learning perspective with a parameterized policy that we theoretically prove to act as a dynamic linear surrogate maximizing the constructive interference of gradients. To enhance practical flexibility, AC-ODM supports two operational modes: (i) a proxy mode for fixed, pre-prepared corpora, where a policy learned on a small model is transferred to a larger target; and (ii) a non-proxy mode for direct end-to-end training from scratch without priors. Empirically, AC-ODM significantly outperforms prior methods in convergence speed and downstream accuracy across various architectures. On Pythia-1B, it reaches optimal validation perplexity using up to 66% fewer training steps than competitive baselines, delivering a 27.5% relative improvement in MMLU accuracy and a 2.23 x higher pass@1 on HumanEval, all while incurring a virtually negligible (0.4%) per-step wall-clock increase and only 2% additional memory overhead. Code is available at https://github.com/DANG-ai/AC-ODM.

cs.LG

MeshAC: A 3D Mesh Generation and Adaptation Package for Multiscale Coupling Methods

This paper introduces the MeshAC package, which generates three-dimensional adaptive meshes tailored for the efficient and robust implementation of multiscale coupling methods. While Delaunay triangulation is commonly used for mesh generation across the entire computational domain, generating meshes for multiscale coupling methods is more challenging due to intrinsic discrete structures such as defects, and the need to match these structures to the continuum domain at the interface. The MeshAC package tackles these challenges by generating meshes that align with fine-level discrete structures. It also incorporates localized modification and reconstruction operations specifically designed for interfaces. These enhancements improve both the implementation efficiency and the quality of the coupled mesh. Furthermore, MeshAC introduces a novel adaptive feature that utilizes gradient-based a posteriori error estimation, which automatically adjusts the atomistic region and continuum mesh, ensuring an optimal balance between accuracy and efficiency. This package can be directly applied to the geometry optimization problems of a/c coupling in static mechanics, with potential extensions to many other scenarios. Its capabilities are demonstrated for complex material defects, including straight edge dislocation in BCC W and double voids in FCC Cu. These results suggest that MeshAC can be a valuable tool for researchers and practitioners in computational mechanics.

cs.GR

Adaptive Multigrid Strategy for Geometry Optimization of Large-Scale Three Dimensional Molecular Mechanics

In this paper, we present an efficient adaptive multigrid strategy for the geometry optimization of large-scale three dimensional molecular mechanics. The resulting method can achieve significantly reduced complexity by exploiting the intrinsic low-rank property of the material configurations and by combining the state-of-the-art adaptive techniques with the hierarchical structure of multigrid algorithms. To be more precise, we develop a oneway multigrid method with adaptive atomistic/continuum (a/c) coupling, e.g., blended ghost force correction (BGFC) approximations with gradient-based a posteriori error estimators on the coarse levels. We utilize state-of-the-art 3D mesh generation techniques to effectively implement the method. For 3D crystalline defects, such as vacancies, micro-cracks and dislocations, compared with brute-force optimization, complexity with superior rates can be observed numerically, and the strategy has a five-fold acceleration in terms of CPU time for systems with $10^8$ atoms.

physics.comp-ph

A Posteriori Error Estimates for Adaptive QM/MM Coupling Methods

Hybrid quantum/molecular mechanics models (QM/MM methods) are widely used in material and molecular simulations when MM models do not provide sufficient accuracy but pure QM models are computationally prohibitive. Adaptive QM/MM coupling methods feature on-the-fly classification of atoms during the simulation, allowing the QM and MM subsystems to be updated as needed. In this work, we propose such an adaptive QM/MM method for material defect simulations based on a new residual based it a posteriori error estimator, which provides both lower and upper bounds for the true error. We validate the analysis and illustrate the effectiveness of the new scheme on numerical simulations for material defects.

math.NA

A consistency study of coarse-grained dynamical chains through a nonlinear wave equation of mixed type

A dynamical atomistic chain to simulate mechanical properties of a one-dimensional material with zero temperature may be modelled by the molecular dynamics (MD) model. Because the number of particles (atoms) is huge for a MD model, in practice one often takes a much smaller number of particles to formulate a coarse-grained approximation. We shall mainly consider the consistency of the coarse-grained model with respect to the grain (mesh) size to provide a justification to the goodness of such an approximation. In order to reduce the characteristic oscillations with very different frequencies in such a model, we either add a viscous term to the coarse-grained MD model or apply a space average to the coarse-grained MD solutions for the consistency study. The coarse-grained solution is also compared with the solution of the (macroscopic) continuum model (a nonlinear wave equation of mixed type) to show how well the coarse-grained model can approximate the macroscopic behavior of the material. We also briefly study the instability of the dynamical atomistic chain and the solution of the Riemann problem of the continuum model which may be related to the defect of the atomistic chain under a large deformation in certain locations.

math.NA

Adaptive QM/MM Coupling for Crystalline Defects

QM (quantum mechenics) and MM (molecular mechenics) coupling methods are widely used in simulations of crystalline defects. In this paper, we construct a residual based a posteriori error indicator for QM/MM coupling approximations. We prove the reliability of the error indicator (upper bound of the true approximation error) and develop some sampling techniques for its efficient calculation. Based on the error indicator and Dörfler marking strategy, we design an adaptive QM/MM algorithm for crystalline defects and demonstrate the efficiency with some numerical experiments.

math.NA

A Posteriori Error Estimate and Adaptive Mesh Refinement Algorithm for Atomistic/Continuum Coupling with Finite Range Interactions in Two Dimensions

In this paper, we develop the residual based a posteriori error estimates and the corresponding adaptive mesh refinement algorithm for atomistic/continuum (a/c) coupling with finite range interactions in two dimensions. We have systematically derived a new explicitly computable stress tensor formula for finite range interactions. In particular, we use the geometric reconstruction based consistent atomistic/continuum (GRAC) coupling scheme, which is optimal if the continuum model is discretized by $P^1$ finite elements. The numerical results of the adaptive mesh refinement algorithm is consistent with the optimal a priori error estimates.

math.NA

A Posteriori Error Estimation and Adaptive Algorithm for the Atomistic/Continuum Coupling in 2D

Atomistic/continuum coupling methods aim to achieve optimal balance between accuracy and efficiency. Adaptivity is the key for the efficient implementation of such methods. In this paper, we carry out a rigorous a posteriori analysis of the residual, the stability constant, and the error bound, for a consistent atomistic/continuum coupling method in 2D. We design and implement the corresponding adaptive mesh refinement algorithm, and the convergence rate with respect to degrees of freedom is optimal compare with a priori error estimates.

math.NA