SearcharxivSearch

arXiv subjects

Yuan Pei

Publications and source records attributed to Yuan Pei.

13 recordsLinked to original sources

CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM

Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational heterogeneity. Current solvers are often limited to static predictions or require computationally intensive heuristic searches. We present CryoACE, an end-to-end framework that reconstructs precise atomic graphs for both homogeneous and heterogeneous structures. Our method features two key innovations: an atom-centric reconstruction paradigm, where density features are sampled directly at atomic coordinates and iteratively recycled to refine structures, replacing expensive voxel convolutions for efficient multimodal fusion; and a training-free guidance mechanism that leverages predicted local resolution priors to resolve dynamic ambiguity. Validated on a newly constructed high-quality dataset, CryoACE significantly outperforms existing baselines on static benchmarks and, for the first time, unveils atomic-level dynamic conformations on complex real-world datasets like EMPIAR-10345 without relying on pre-built static structures.

cs.AI

On a Partial Voigt Regularization of the 3D Magnetohydrodynamic Equations in Velocity-Vorticity Form

The Velocity-Vorticity (VV) formulation of the incompressible Navier-Stokes equations has become popular in recent years, especially in numerical studies, due to its structural advantages. Recently, with L. Rebholz, we introduced a Voigt regularization to the momentum equation in this formulation, establishing global well-posedness of the regularized system in 3D, along with convergence results and a blow-up criterion. In the present work, we extend these ideas to the 3D magnetohydrodynamics (MHD) equations. While it may seem that a ``VV-type'' split on the magnetic equation is required, we show that no such modification is necessary, and global well-posedness holds with a Voigt regularization only on the momentum equation, preserving the structure of both the vorticity and magnetic equations. We also prove that the regularized system converges to the original system, up to a possible blow-up time, and we establish a blow-up criterion for solutions to the original 3D MHD system.

math.AP

CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy

Pose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in scientific imaging fields like cryo-electron microscopy (cryo-EM) for near-atomic protein reconstruction. In cryo-EM, pose estimation and 3D reconstruction from unordered particle images still depend on time-consuming iterative optimization, primarily due to challenges such as low signal-to-noise ratios (SNR) and distortions from the contrast transfer function (CTF). We introduce CryoFastAR, the first geometric foundation model that can directly predict poses from Cryo-EM noisy images for Fast ab initio Reconstruction. By integrating multi-view features and training on large-scale simulated cryo-EM data with realistic noise and CTF modulations, CryoFastAR enhances pose estimation accuracy and generalization. To enhance training stability, we propose a progressive training strategy that first allows the model to extract essential features under simpler conditions before gradually increasing difficulty to improve robustness. Experiments show that CryoFastAR achieves comparable quality while significantly accelerating inference over traditional iterative approaches on both synthetic and real datasets.

cs.CV

DRACO: A Denoising-Reconstruction Autoencoder for Cryo-EM

Foundation models in computer vision have demonstrated exceptional performance in zero-shot and few-shot tasks by extracting multi-purpose features from large-scale datasets through self-supervised pre-training methods. However, these models often overlook the severe corruption in cryogenic electron microscopy (cryo-EM) images by high-level noises. We introduce DRACO, a Denoising-Reconstruction Autoencoder for CryO-EM, inspired by the Noise2Noise (N2N) approach. By processing cryo-EM movies into odd and even images and treating them as independent noisy observations, we apply a denoising-reconstruction hybrid training scheme. We mask both images to create denoising and reconstruction tasks. For DRACO's pre-training, the quality of the dataset is essential, we hence build a high-quality, diverse dataset from an uncurated public database, including over 270,000 movies or micrographs. After pre-training, DRACO naturally serves as a generalizable cryo-EM image denoiser and a foundation model for various cryo-EM downstream tasks. DRACO demonstrates the best performance in denoising, micrograph curation, and particle picking tasks compared to state-of-the-art baselines.

cs.CV

Calmed Ohmic Heating for the 2D Magnetohydrodynamic-Boussinesq System: Global Well-posedness and Convergence

When an electric current runs through a fluid, it generates heat via a process known as ``Ohmic heating'' or ``Joule heating.'' While this phenomenon, and its quantification known as Joule's Law, is the first studied example of heat generation via an electric field, many difficulties still remain in understanding its consequences. In particular, a magnetic fluid naturally generates an electric field via Ampère's law, which heats the fluid via Joule's law. This heat in turn gives rise to convective effects in the fluid, creating complicated dynamical behavior. This has been modeled (in other works) by including an Ohmic heating term in the Magnetohydrodynamic-Boussinessq (MHD-B) equation. However, the structure of this term causes major analytical difficulties, and basic questions of well-posedness remain open problems, even in the two-dimensional case. Moreover, standard approaches to finding a globally well-posed approximate model, such as filtering or adding high-order diffusion, are not enough to handle the Ohmic heating term. In this work, we present a different approach that we call ``calming'', which reduces the effective algebraic degree of the Ohmic heating term in a controlled manner. We show that this new model is globally well-posed, and moreover, its solutions converge to solutions of the MHD-B system with the Ohmic heating term (assuming that solutions to the original equation exist), making it the first globally well-posed approximate model for the MHD-B equation with Ohmic heating.

math.AP

Breaking MLPerf Training: A Case Study on Optimizing BERT

Speeding up the large-scale distributed training is challenging in that it requires improving various components of training including load balancing, communication, optimizers, etc. We present novel approaches for fast large-scale training of BERT model which individually ameliorates each component thereby leading to a new level of BERT training performance. Load balancing is imperative in distributed BERT training since its training datasets are characterized by samples with various lengths. Communication cost, which is proportional to the scale of distributed training, needs to be hidden by useful computation. In addition, the optimizers, e.g., ADAM, LAMB, etc., need to be carefully re-evaluated in the context of large-scale distributed training. We propose two new ideas, (1) local presorting based on dataset stratification for load balancing and (2) bucket-wise gradient clipping before allreduce which allows us to benefit from the overlap of gradient computation and synchronization as well as the fast training of gradient clipping before allreduce. We also re-evaluate existing optimizers via hyperparameter optimization and utilize ADAM, which also contributes to fast training via larger batches than existing methods. Our proposed methods, all combined, give the fastest MLPerf BERT training of 25.1 (22.3) seconds on 1,024 NVIDIA A100 GPUs, which is 1.33x (1.13x) and 1.57x faster than the other top two (one) submissions to MLPerf v1.1 (v2.0). Our implementation and evaluation results are available at MLPerf v1.1~v2.1.

cs.LG

The second-best way to do sparse-in-time continuous data assimilation: Improving convergence rates for the 2D and 3D Navier-Stokes equations

We study different approaches to implementing sparse-in-time observations into the the Azouani-Olson-Titi data assimilation algorithm. We propose a new method which introduces a "data assimilation window" separate from the observational time interval. We show that by making this window as small as possible, we can drastically increase the strength of the nudging parameter without losing stability. Previous methods used old data to nudge the solution until a new observation was made. In contrast, our method stops nudging the system almost immediately after an observation is made, allowing the system relax to the correct physics. We show that this leads to an order-of-magnitude improvement in the time to convergence in our 3D Navier-Stokes simulations. Moreover, our simulations indicate that our approach converges at nearly the same rate as the idealized method of direct replacement of low Fourier modes proposed by Hayden, Olson, and Titi (HOT). However, our approach can be readily adapted to non-idealized settings, such as finite element methods, finite difference methods, etc., since there is no need to access Fourier modes as our method works for general interpolants. It is in this sense that we think of our approach as ``second best;'' that is, the ``best'' method would be the direct replacement of Fourier modes as in HOT, but this idealized approach is typically not feasible in physically realistic settings. While our method has a convergence rate that is slightly sub-optimal compared to the idealized method, it is directly compatible with real-world applications. Moreover, we prove analytically that these new algorithms are globally well-posed, and converge to the true solution exponentially fast in time. In addition, we provide the first 3D computational validation of HOT algorithm.

math.AP

Approximate continuous data assimilation of the 2D Navier-Stokes equations via the Voigt-regularization with observable data

We propose a data assimilation algorithm for the 2D Navier-Stokes equations, based on the Azouani, Olson, and Titi (AOT) algorithm, but applied to the 2D Navier-Stokes-Voigt equations. Adapting the AOT algorithm to regularized versions of Navier-Stokes has been done before, but the innovation of this work is to drive the assimilation equation with observational data, rather than data from a regularized system. We first prove that this new system is globally well-posed. Moreover, we prove that for any admissible initial data, the $L^2$ and $H^1$ norms of error are bounded by a constant times a power of the Voigt-regularization parameter $α>0$, plus a term which decays exponentially fast in time. In particular, the large-time error goes to zero algebraically as $α$ goes to zero. Assuming more smoothness on the initial data and forcing, we also prove similar results for the $H^2$ norm.

math.AP

Continuous data assimilation for the 3D primitive equations of the ocean

In this article, we show that the continuous data assimilation algorithm is valid for the 3D primitive equations of the ocean. Namely, the $L^2$ norm of the assimilated solution converge to that of the reference solution at an exponential rate in time. We also prove the global existence of strong solution to the assimilated system.

math.AP

Global well-posedness of the velocity-vorticity-Voigt model of the 3D Navier-Stokes equations

The velocity-vorticity formulation of the 3D Navier-Stokes equations was recently found to give excellent numerical results for flows with strong rotation. In this work, we propose a new regularization of the 3D Navier-Stokes equations, which we call the 3D velocity-vorticity-Voigt (VVV) model, with a Voigt regularization term added to momentum equation in velocity-vorticity form, but with no regularizing term in the vorticity equation. We prove global well-posedness and regularity of this model under periodic boundary conditions. We prove convergence of the model's velocity and vorticity to their counterparts in the 3D Navier-Stokes equations as the Voigt modeling parameter tends to zero. We prove that the curl of the model's velocity converges to the model vorticity (which is solved for directly), as the Voigt modeling parameter tends to zero. Finally, we provide a criterion for finite-time blow-up of the 3D Navier-Stokes equations based on this inviscid regularization.

math.AP

Continuous data assimilation for the magnetohydrodynamic equations in 2D using one component of the velocity and magnetic fields

We propose several continuous data assimilation (downscaling) algorithms based on feedback control for the 2D magnetohydrodynamic (MHD) equations. We show that for sufficiently large choices of the control parameter and resolution and assuming that the observed data is error-free, the solution of the controlled system converges exponentially (in $L^2$ and $H^1$ norms) to the reference solution independently of the initial data chosen for the controlled system. Furthermore, we show that a similar result holds when controls are placed only on the horizontal (or vertical) variables, or on a single Elsässer variable, under more restrictive conditions on the control parameter and resolution. Finally, using the data assimilation system, we show the existence of abridged determining modes, nodes and volume elements.

math.AP

Nonlinear Continuous Data Assimilation

We introduce three new nonlinear continuous data assimilation algorithms. These models are compared with the linear continuous data assimilation algorithm introduced by Azouani, Olson, and Titi (AOT). As a proof-of-concept for these models, we computationally investigate these algorithms in the context of the 1D Kuramoto-Sivashinsky equation. We observe that the nonlinear models experience super-exponential convergence in time, and converge to machine precision significantly faster than the linear AOT algorithm in our tests.

math.AP

On the local well-posedness and a Prodi-Serrin type regularity criterion of the three-dimensional MHD-Boussinesq system without thermal diffusion

We prove a Prodi-Serrin-type global regularity condition for the three-dimensional Magnetohydrodynamic-Boussinesq system (3D MHD-Boussinesq) without thermal diffusion, in terms of only two velocity and two magnetic components. This is the first Prodi-Serrin-type criterion for a hydrodynamic system which is not fully dissipative, and indicates that such an approach may be successful on other systems. In addition, we provide a constructive proof of the local well-posedness of solutions to the fully dissipative 3D MHD-Boussinesq system, and also the fully inviscid, irresistive, non-diffusive MHD-Boussinesq equations. We note that, as a special case, these results include the 3D non-diffusive Boussinesq system and the 3D MHD equations. Moreover, they can be extended without difficulty to include the case of a Coriolis rotational term.

math.AP