SearcharxivSearch

arXiv subjects

Dongdong He

Publications and source records attributed to Dongdong He.

18 recordsLinked to original sources

Training Report of TeleChat3-MoE

TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one trillion,trained end-to-end on Ascend NPU cluster. This technical report mainly presents the underlying training infrastructure that enables reliable and efficient scaling to frontier model sizes. We detail systematic methodologies for operator-level and end-to-end numerical accuracy verification, ensuring consistency across hardware platforms and distributed parallelism strategies. Furthermore, we introduce a suite of performance optimizations, including interleaved pipeline scheduling, attention-aware data scheduling for long-sequence training,hierarchical and overlapped communication for expert parallelism, and DVM-based operator fusion. A systematic parallelization framework, leveraging analytical estimation and integer linear programming, is also proposed to optimize multi-dimensional parallelism configurations. Additionally, we present methodological approaches to cluster-level optimizations, addressing host- and device-bound bottlenecks during large-scale training tasks. These infrastructure advancements yield significant throughput improvements and near-linear scaling on clusters comprising thousands of devices, providing a robust foundation for large-scale language model development on hardware ecosystems.

cs.CL

Linear, decoupled, second-order and structure-preserving scheme for Carreau fluid equations coupled with steric Poisson-Nernst-Planck model

In this paper, to study ionic steric effects, we present a linear, decoupled, second-order accurate in time and structure-preserving scheme with finite element approximations for Carreau fluid equations coupled with steric Poisson-Nernst-Planck (SPNP) model. The logarithmic transformation for the ion concentration is used to preserve positivity property. To deal with the nonlinear coupling terms in fluid equation, a nonlocal auxiliary variable with respect to the free energy of SPNP equations and its associated ordinary differential equation are introduced. The obtained system is equivalent to the original system. The fully discrete scheme is proved to be mass conservative, positivity-preserving for ion concentration and energy dissipative at discrete level. Some numerical simulations are provided to demonstrate its stability and accuracy. Moreover, the ionic steric effects are numerically investigated.

math.NA

Linear, decoupled, positivity preserving, positive-definiteness preserving and energy stable schemes for the diffusive Oldroyd-B coupled with PNP model

In this paper, we present a first-order finite element scheme for the viscoelastic electrohydrodynamic model. The model incorporates the Poisson-Nernst-Planck equations to describe the transport of ions and the Oldroyd-B constitutive model to capture the behavior of viscoelastic fluids. To preserve the positive-definiteness of the conformation tensor and the positivity of ion concentrations, we employ both logarithmic transformations. The decoupled scheme is achieved by introducing a nonlocal auxiliary variable and using the splitting technique. The proposed schemes are rigorously proven to be mass conservative and energy stable at the fully discrete level. To validate the theoretical analysis, we present numerical examples that demonstrate the convergence rates and the robust performance of the schemes. The results confirm that the proposed methods accurately handle the high Weissenberg number problem (HWNP) at moderately high Weissenberg numbers. Finally, the flow structure influenced by the elastic effect within the electro-convection phenomena has been studied.

math.NA

Monocular Depth Guided Occlusion-Aware Disparity Refinement via Semi-supervised Learning in Laparoscopic Images

Occlusion and the scarcity of labeled surgical data are significant challenges in disparity estimation for stereo laparoscopic images. To address these issues, this study proposes a Depth Guided Occlusion-Aware Disparity Refinement Network (DGORNet), which refines disparity maps by leveraging monocular depth information unaffected by occlusion. A Position Embedding (PE) module is introduced to provide explicit spatial context, enhancing the network's ability to localize and refine features. Furthermore, we introduce an Optical Flow Difference Loss (OFDLoss) for unlabeled data, leveraging temporal continuity across video frames to improve robustness in dynamic surgical scenes. Experiments on the SCARED dataset demonstrate that DGORNet outperforms state-of-the-art methods in terms of End-Point Error (EPE) and Root Mean Squared Error (RMSE), particularly in occlusion and texture-less regions. Ablation studies confirm the contributions of the Position Embedding and Optical Flow Difference Loss, highlighting their roles in improving spatial and temporal consistency. These results underscore DGORNet's effectiveness in enhancing disparity estimation for laparoscopic surgery, offering a practical solution to challenges in disparity estimation and data limitations.

cs.CV

Drawing of Weakly Viscoelastic Fluid Tubes

We explore the drawing of an axisymmetric viscoelastic tube subject to inertial and surface tension effects. We adopt the Giesekus constitutive model and derive asymptotic long-wave equations for weakly viscoelastic effects. Intuitively, one might imagine that the elastic stresses should act to prevent hole closure during the drawing process. Surprisingly, our results show that the hole closure at the outlet is enhanced by elastic effects for most parameter values. However, the opposite is true if the tube has a very large hole size at the inlet of the device or if the axial stretching is very weak. We explain the physical mechanism underlying this phenomenon by examining how the second normal stress difference induced by elastic effects modifies the hole evolution process. We also determine how viscoelasticity affects the stability of the drawing process and show that elastic effects are always destabilizing for negligible inertia. This is in direct contrast to the case of a thread without a hole for which elastic effects are always stabilizing. On the other hand, our results show that if the inertia is non-zero, elastic effects can be either stabilizing or destabilizing depending on the parameters.

physics.flu-dyn

Decomposition approach for Stackelberg P-median problem with user preferences

The P-median facility location problem with user preferences (PUP) studies an operator that locates P facilities to serve customers/users in a cost-efficient manner, upon anticipating customer preferences and choices. The problem can be visualized as a leader-follower game in which the operator is the leader that opens facilities, whereas the customer is the follower who observes the operator's location decision at first and then seeks services from the most preferred facility. Such a modeling perspective is of practical importance as we have witnessed its applications to various problems, such as the establishment of power plants in energy markets and the location of healthcare service centers for COVID-19 Vaccination. Despite that a considerable number of solution methodologies have been proposed, many of them are heuristic methods whose solution quality cannot be easily verified. Moreover, due to the hardness of the problems, existing exact approaches have limited performance. Motivated by these observations, we aim to develop an efficient exact algorithm for solving large-scale PUP models. We first propose a branch-and-cut decomposition algorithm and then design accelerated techniques to further enhance the performance. Using a broad testbed, we show that our algorithm outperforms various exact approaches by a large margin, and the advantage can go up to several orders of magnitude in terms of computational time in some datasets. Finally, we conduct sensitivity analysis to draw additional implications and to highlight the importance of considering user preferences when they exist.

math.OC

A conservative difference scheme with optimal pointwise error estimates for two-dimensional space fractional nonlinear Schrödinger equations

In this paper, a linearized semi-implicit finite difference scheme is proposed for solving the two-dimensional (2D) space fractional nonlinear Schrödinger equation (SFNSE).The scheme has the property of mass and energy conservation on the discrete level, with an unconditional stability and a second order accuracy for both time and spatial variables. The main contribution of this paper is an optimal pointwise error estimate for the 2D SFNSE, which is rigorously established and proved for the first time. Moreover, a novel technique is proposed for dealing with the nonlinear term in the equation, which plays an essential role in the error estimation. Finally, the numerical results confirm well with the theoretical findings.

math.NA

Last-mile Delivery: Optimal Locker Location Under Multinomial Logit Choice Model

One innovative solution to the last-mile delivery problem is the self-service locker system. Motivated by a real case in Singapore, we consider a POP-Locker Alliance who operates a set of POP-stations and wishes to improve the last-mile delivery by opening new locker facilities. We propose a quantitative approach to determine the optimal locker location with the objective to maximize the overall service provided by the alliance. Customer's choices regarding the use of facilities are explicitly considered. They are predicted by a multinomial logit model. We then formulate the location problem as a multi-ratio linear-fractional 0-1 program and provide two solution approaches. The first one is to reformulate the original problem as a mixed-integer linear program, which is further strengthened using conditional McCormick inequalities. This approach is an exact method, developed for small-scale problems. For large-scale problems, we propose a Suggest-and-Improve framework with two embedded algorithms. Numerical studies indicated that our framework is an efficient approach that yields high-quality solutions. Finally, we conducted a case study. The results highlighted the importance of considering the customers' choices. Under different parameter values of the multinomial logit model, the decisions could be completely different. Therefore, the parameter value should be carefully estimated in advance.

math.OC

A linearized energy--conservative finite element method for the nonlinear Schr\"{o}dinger equation with wave operator

In this paper, we propose a linearized finite element method (FEM) for solving the cubic nonlinear Schr\"{o}dinger equation with wave operator. In this method, a modified leap-frog scheme is applied for time discretization and a Galerkin finite element method is applied for spatial discretization. We prove that the proposed method keeps the energy conservation in the given discrete norm. Comparing with non-conservative schemes, our algorithm keeps higher stability. Meanwhile, an optimal error estimate for the proposed scheme is given by an error splitting technique. That is, we split the error into two parts, one from temporal discretization and the other from spatial discretization. First, by introducing a time-discrete system, we prove the uniform boundedness for the solution of this time-discrete system in some strong norms and obtain error estimates in temporal direction. With the help of the preliminary temporal estimates, we then prove the pointwise uniform boundedness of the finite element solution, and obtain the optimal $L^2$-norm error estimates in the sense that the time step size is not related to spatial mesh size. Finally, numerical examples are provided to validate the convergence-order, unconditional stability and energy conservation.

math.NA

An efficient multigrid solver for 3D biharmonic equation with a discretization by 25-point difference scheme

In this paper, we propose an efficient extrapolation cascadic multigrid (EXCMG) method combined with 25-point difference approximation to solve the three-dimensional biharmonic equation. First, through applying Richardson extrapolation and quadratic interpolation on numerical solutions on current and previous grids, a third-order approximation to the finite difference solution can be obtained and used as the iterative initial guess on the next finer grid. Then we adopt the bi-conjugate gradient (Bi-CG) method to solve the large linear system resulting from the 25-point difference approximation. In addition, an extrapolation method based on midpoint extrapolation formula is used to achieve higher-order accuracy on the entire finest grid. Finally, some numerical experiments are performed to show that the EXCMG method is an efficient solver for the 3D biharmonic equation.

math.NA

A three-level linearized difference scheme for the coupled nonlinear fractional Ginzburg-Landau equation

In this paper, the coupled fractional Ginzburg-Landau equations are first time investigated numerically. A linearized implicit finite difference scheme is proposed. The scheme involves three time levels, is unconditionally stable and second-order accurate in both time and space variables. The unique solvability, the unconditional stability and optimal pointwise error estimates are obtained by using the energy method and mathematical induction. Moreover, the proposed second-order method can be easily extended into the fourth-order method by using an average finite difference operator for spatial fractional derivatives and Richardson extrapolation for time variable. Finally, numerical results are presented to confirm the theoretical results.

math.NA

A fourth-order maximum principle preserving operator splitting scheme for three-dimensional fractional Allen-Cahn equations

In this paper, by using Strang's second-order splitting method, the numerical procedure for the three-dimensional (3D) space fractional Allen-Cahn equation can be divided into three steps. The first and third steps involve an ordinary differential equation, which can be solved analytically. The intermediate step involves a 3D linear fractional diffusion equation, which is solved by the Crank-Nicolson alternating directional implicit (ADI) method. The ADI technique can convert the multidimensional problem into a series of one-dimensional problems, which greatly reduces the computational cost. A fourth-order difference scheme is adopted for discretization of the space fractional derivatives. Finally, Richardson extrapolation is exploited to increase the temporal accuracy. The proposed method is shown to be unconditionally stable by Fourier analysis. Another contribution of this paper is to show that the numerical solutions satisfy the discrete maximum principle under reasonable time step constraint. For fabricated smooth solutions, numerical results show that the proposed method is unconditionally stable and fourth-order accurate in both time and space variables. In addition, the discrete maximum principle is also numerically verified.

math.NA

Wave analysis in one dimensional structures with a wavelet finite element model and precise integration method

Numerical simulation of ultrasonic wave propagation provides an efficient tool for crack identification in structures, while it requires a high resolution and expensive time calculation cost in both time integration and spatial discretization. Wavelet finite element model provides a highorder finite element model and gives a higher accuracy on spatial discretization, B-Spline wavelet interval (BSWI) has been proved to be one of the most commonly used wavelet finite element model with the advantage of getting the same accuracy but with fewer element so that the calculation cost is much lower than traditional finite element method and other high-order element methods. Precise Integration Method provides a higher resolution in time integration and has been proved to be a stable time integration method with a much lower cut-off error for same and even smaller time step. In this paper, a wavelet finite element model combined with precise integration method is presented for the numerical simulation of ultrasonic wave propagation and crack identification in 1D structures. Firstly, the wavelet finite element based on BSWI is constructed for rod and beam structures. Then Precise Integrated Method is introduced with application for the wave propagation in 1D structures. Finally, numerical examples of ultrasonic wave propagation in rod and beam structures are conducted for verification. Moreover, crack identification in both rod and beam structures are studied based on the new model.

cs.CE

A linearly implicit conservative difference scheme for the generalized Rosenau-Kawahara-RLW equation

This paper concerns the numerical study for the generalized Rosenau-Kawahara-RLW equation obtained by coupling the generalized Rosenau-RLW equation and the generalized Rosenau-Kawahara equation. We first derive the energy conservation law of the equation, and then develop a three-level linearly implicit difference scheme for solving the equation. We prove that the proposed scheme is energy-conserved, unconditionally stable and second-order accurate both in time and space variables. Finally, numerical experiments are carried out to confirm the energy conservation, the convergence rates of the scheme and effectiveness for long-time simulation.

math.NA

An extrapolation cascadic multigrid method combined with a fourth order compact scheme for 3D poisson equation

In this paper, we develop an EXCMG method to solve the three-dimensional Poisson equation on rectangular domains by using the compact finite difference (FD) method with unequal meshsizes in different coordinate directions. The resulting linear system from compact FD discretization is solved by the conjugate gradient (CG) method with a relative residual stopping criterion. By combining the Richardson extrapolation and tri-quartic Lagrange interpolation for the numerical solutions from two-level of grids (current and previous grids), we are able to produce an extremely accurate approximation of the actual numerical solution on the next finer grid, which can greatly reduce the number of relaxation sweeps needed. Additionally, a simple method based on the midpoint extrapolation formula is used for the fourth-order FD solutions on two-level of grids to achieve sixth-order accuracy on the entire fine grid cheaply and directly. The gradient of the numerical solution can also be easily obtained through solving a series of tridiagonal linear systems resulting from the fourth-order compact FD discretizations. Numerical results show that our EXCMG method is much more efficient than the classical V-cycle and W-cycle multigrid methods. Moreover, only few CG iterations are required on the finest grid to achieve full fourth-order accuracy in both the $L^2$-norm and $L^{\infty}$-norm for the solution and its gradient when the exact solution belongs to $C^6$. Finally, numerical result shows that our EXCMG method is still effective when the exact solution has a lower regularity, which widens the scope of applicability of our EXCMG method.

math.NA

A new extrapolation cascadic multigrid method for 3D elliptic boundary value problems on rectangular domains

In this paper, we develop a new extrapolation cascadic multigrid (ECMG$_{jcg}$) method, which makes it possible to solve 3D elliptic boundary value problems on rectangular domains of over 100 million unknowns on a desktop computer in minutes. First, by combining Richardson extrapolation and tri-quadratic Serendipity interpolation techniques, we introduce a new extrapolation formula to provide a good initial guess for the iterative solution on the next finer grid, which is a third order approximation to the finite element (FE) solution. And the resulting large sparse linear system from the FE discretization is then solved by the Jacobi-preconditioned Conjugate Gradient (JCG) method. Additionally, instead of performing a fixed number of iterations as cascadic multigrid (CMG) methods, a relative residual stopping criterion is used in iterative solvers, which enables us to obtain conveniently the numerical solution with the desired accuracy. Moreover, a simple Richardson extrapolation is used to cheaply get a fourth order approximate solution on the entire fine grid. Test results are reported to show that ECMG$_{jcg}$ has much better efficiency compared to the classical MG methods. Since the initial guess for the iterative solution is a quite good approximation to the FE solution, numerical results show that only few number of iterations are required on the finest grid for ECMG$_{jcg}$ with an appropriate tolerance of the relative residual to achieve full second order accuracy, which is particularly important when solving large systems of equations and can greatly reduce the computational cost. It should be pointed out that when the tolerance becomes smaller, ECMG$_{jcg}$ still needs only few iterations to obtain fourth order extrapolated solution on each grid, except on the finest grid. Finally, we present the reason why our ECMG algorithms are so highly efficient for solving such problems.

math.NA

An energy preserving finite difference scheme for the Poisson-Nernst-Planck system

In this paper, we construct a semi-implicit finite difference method for the time dependent Poisson-Nernst-Planck system. Although the Poisson-Nernst-Planck system is a nonlinear system, the numerical method presented in this paper only needs to solve a linear system at each time step, which can be done very efficiently. The rigorous proof for the mass conservation and electric potential energy decay are shown. Moreover, mesh refinement analysis shows that the method is second order convergent in space and first order convergent in time. Finally we point out that our method can be easily extended to the case of multi-ions.

math.NA

A mathematical model of the metabolic and perfusion effects on cortical spreading depression

Cortical spreading depression (CSD) is a slow-moving ionic and metabolic disturbance that propagates in cortical brain tissue. In addition to massive cellular depolarization, CSD also involves significant changes in perfusion and metabolism -- aspects of CSD that had not been modeled and are important to traumatic brain injury, subarachnoid hemorrhage, stroke, and migraine. In this study, we develop a mathematical model for CSD where we focus on modeling the features essential to understanding the implications of neurovascular coupling during CSD. In our model, the sodium-potassium--ATPase, mainly responsible for ionic homeostasis and active during CSD, operates at a rate that is dependent on the supply of oxygen. The supply of oxygen is determined by modeling blood flow through a lumped vascular tree with an effective local vessel radius that is controlled by the extracellular potassium concentration. We show that during CSD, the metabolic demands of the cortex exceed the physiological limits placed on oxygen delivery, regardless of vascular constriction or dilation. However, vasoconstriction and vasodilation play important roles in the propagation of CSD and its recovery. Our model replicates the qualitative and quantitative behavior of CSD -- vasoconstriction, oxygen depletion, extracellular potassium elevation, prolonged depolarization -- found in experimental studies. We predict faster, longer duration CSD in vivo than in vitro due to the contribution of the vasculature. Our results also help explain some of the variability of CSD between species and even within the same animal. These results have clinical and translational implications, as they allow for more precise in vitro, in vivo, and in silico exploration of a phenomenon broadly relevant to neurological disease.

q-bio.NC