SearcharxivSearch

arXiv subjects

Ning Tian

Publications and source records attributed to Ning Tian.

10 recordsLinked to original sources

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

cs.CL

DeepSeek-V3 Technical Report

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.

cs.CL

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

We present DeepSeek-V2, a strong Mixture-of-Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token, and supports a context length of 128K tokens. DeepSeek-V2 adopts innovative architectures including Multi-head Latent Attention (MLA) and DeepSeekMoE. MLA guarantees efficient inference through significantly compressing the Key-Value (KV) cache into a latent vector, while DeepSeekMoE enables training strong models at an economical cost through sparse computation. Compared with DeepSeek 67B, DeepSeek-V2 achieves significantly stronger performance, and meanwhile saves 42.5% of training costs, reduces the KV cache by 93.3%, and boosts the maximum generation throughput to 5.76 times. We pretrain DeepSeek-V2 on a high-quality and multi-source corpus consisting of 8.1T tokens, and further perform Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) to fully unlock its potential. Evaluation results show that, even with only 21B activated parameters, DeepSeek-V2 and its chat versions still achieve top-tier performance among open-source models.

cs.CL

Metal to Mott Insulator Transition in Two-dimensional 1T-TaSe$_2$

When electron-electron interaction dominates over other electronic energy scales, exotic, collective phenomena often emerge out of seemingly ordinary matter. The strongly correlated phenomena, such as quantum spin liquid and unconventional superconductivity, represent a major research frontier and a constant source of inspiration. Central to strongly correlated physics is the concept of Mott insulator, from which various other correlated phases derive. The advent of two-dimensional (2D) materials brings unprecedented opportunities to the study of strongly correlated physics in the 2D limit. In particular, the enhanced correlation and extreme tunability of 2D materials enables exploring strongly correlated systems across uncharted parameter space. Here, we discover an intriguing metal to Mott insulator transition in 1T-TaSe$_2$ as the material is thinned down to atomic thicknesses. Specifically, we discover, for the first time, that the bulk metallicity of 1T-TaSe$_2$ arises from a band crossing Fermi level. Reducing the dimensionality effectively quenches the kinetic energy of the initially itinerant electrons and drives the material into a Mott insulating state. The dimensionality-driven Metal to Mott insulator transition resolves the long-standing dichotomy between metallic bulk and insulating surface of 1T-TaSe$_2$. Our results additionally establish 1T-TaSe$_2$ as an ideal variable system for exploring various strongly correlated phenomena.

cond-mat.str-el

Real-Time Optimal Lithium-Ion Battery Charging Based on Explicit Model Predictive Control

The rapidly growing use of lithium-ion batteries across various industries highlights the pressing issue of optimal charging control, as charging plays a crucial role in the health, safety and life of batteries. The literature increasingly adopts model predictive control (MPC) to address this issue, taking advantage of its capability of performing optimization under constraints. However, the computationally complex online constrained optimization intrinsic to MPC often hinders real-time implementation. This paper is thus proposed to develop a framework for real-time charging control based on explicit MPC (eMPC), exploiting its advantage in characterizing an explicit solution to an MPC problem, to enable real-time charging control. The study begins with the formulation of MPC charging based on a nonlinear equivalent circuit model. Then, multi-segment linearization is conducted to the original model, and applying the eMPC design to the obtained linear models leads to a charging control algorithm. The proposed algorithm shifts the constrained optimization to offline by precomputing explicit solutions to the charging problem and expressing the charging law as piecewise affine functions. This drastically reduces not only the online computational costs in the control run but also the difficulty of coding. Extensive numerical simulation and experimental results verify the effectiveness of the proposed eMPC charging control framework and algorithm. The research results can potentially meet the needs for real-time battery management running on embedded hardware.

eess.SY

One-Shot Parameter Identification of the Thevenin's Model for Batteries: Methods and Validation

Parameter estimation is of foundational importance for various model-based battery management tasks, including charging control, state-of-charge estimation and aging assessment. However, it remains a challenging issue as the existing methods generally depend on cumbersome and time-consuming procedures to extract battery parameters from data. Departing from the literature, this paper sets the unique aim of identifying all the parameters offline in a one-shot procedure, including the resistance and capacitance parameters and the parameters in the parameterized function mapping from the state-of-charge to the open-circuit voltage. Considering the well-known Thevenin's battery model, the study begins with the parameter identifiability analysis, showing that all the parameters are locally identifiable. Then, it formulates the parameter identification problem in a prediction-error-minimization framework. As the non-convexity intrinsic to the problem may lead to physically meaningless estimates, two methods are developed to overcome this issue. The first one is to constrain the parameter search within a reasonable space by setting parameter bounds, and the other adopts regularization of the cost function using prior parameter guess. The proposed identifiability analysis and identification methods are extensively validated through simulations and experiments.

eess.SY

Nonlinear Double-Capacitor Model for Rechargeable Batteries: Modeling, Identification and Validation

This paper proposes a new equivalent circuit model for rechargeable batteries by modifying a double-capacitor model proposed in [1]. It is known that the original model can address the rate capacity effect and energy recovery effect inherent to batteries better than other models. However, it is a purely linear model and includes no representation of a battery's nonlinear phenomena. Hence, this work transforms the original model by introducing a nonlinear-mapping-based voltage source and a serial RC circuit. The modification is justified by an analogy with the single-particle model. Two parameter estimation approaches, termed 1.0 and 2.0, are designed for the new model to deal with the scenarios of constant-current and variable-current charging/discharging, respectively. In particular, the 2.0 approach proposes the notion of Wiener system identification based on maximum a posteriori estimation, which allows all the parameters to be estimated in one shot while overcoming the nonconvexity or local minima issue to obtain physically reasonable estimates. An extensive experimental evaluation shows that the proposed model offers excellent accuracy and predictive capability. A comparison against the Rint and Thevenin models further points to its superiority. With high fidelity and low mathematical complexity, this model is beneficial for various real-time battery management applications.

eess.SY

Nonlinear Bayesian Estimation: From Kalman Filtering to a Broader Horizon

This article presents an up-to-date tutorial review of nonlinear Bayesian estimation. State estimation for nonlinear systems has been a challenge encountered in a wide range of engineering fields, attracting decades of research effort. To date, one of the most promising and popular approaches is to view and address the problem from a Bayesian probabilistic perspective, which enables estimation of the unknown state variables by tracking their probabilistic distribution or statistics (e.g., mean and covariance) conditioned on the system's measurement data. This article offers a systematic introduction of the Bayesian state estimation framework and reviews various Kalman filtering (KF) techniques, progressively from the standard KF for linear systems to extended KF, unscented KF and ensemble KF for nonlinear systems. It also overviews other prominent or emerging Bayesian estimation methods including the Gaussian filtering, Gaussian-sum filtering, particle filtering and moving horizon estimation and extends the discussion of state estimation forward to more complicated problems such as simultaneous state and parameter/input estimation.

eess.SY

Three-dimensional Temperature Field Reconstruction for A Lithium-Ion Battery Pack: A Distributed Kalman Filtering Approach

Despite the ever-increasing use across different sectors, the lithium-ion batteries (LiBs) have continually seen serious concerns over their thermal vulnerability. The LiB operation is associated with the heat generation and buildup effect, which manifests itself more strongly, in the form of highly uneven thermal distribution, for a LiB pack consisting of multiple cells. If not well monitored and managed, the heating may accelerate aging and cause unwanted side reactions. In extreme cases, it will even cause fires and explosions, as evidenced by a series of well-publicized incidents in recent years. To address this threat, this paper, for the first time, seeks to reconstruct the three-dimensional temperature field of a LiB pack in real time. The major challenge lies in how to acquire a high-fidelity reconstruction with constrained computation time. In this study, a three-dimensional thermal model is established first for a LiB pack configured in series. Although spatially resolved, this model captures spatial thermal behavior with a combination of high integrity and low complexity. Given the model, the standard Kalman filter is then distributed to attain temperature field estimation at substantially reduced computational complexity. The arithmetic operation analysis and numerical simulation illustrate that the proposed distributed estimation achieves a comparable accuracy as the centralized approach but with much less computation. This work can potentially contribute to the safer operation of the LiB packs in various systems dependent on LiB-based energy storage, potentially widening the access of this technology to a broader range of engineering areas.

physics.app-ph

Fast and guaranteed blind multichannel deconvolution under a bilinear system model

We consider the multichannel blind deconvolution problem where we observe the output of multiple channels that are all excited with the same unknown input. From these observations, we wish to estimate the impulse responses of each of the channels. We show that this problem is well-posed if the channels follow a bilinear model where the ensemble of channel responses is modeled as lying in a low-dimensional subspace but with each channel modulated by an independent gain. Under this model, we show how the channel estimates can be found by minimizing a quadratic functional over a non-convex set. We analyze two methods for solving this non-convex program, and provide performance guarantees for each. The first is a method of alternating eigenvectors that breaks the program down into a series of eigenvalue problems. The second is a truncated power iteration, which can roughly be interpreted as a method for finding the largest eigenvector of a symmetric matrix with the additional constraint that it adheres to our bilinear model. As with most non-convex optimization algorithms, the performance of both of these algorithms is highly dependent on having a good starting point. We show how such a starting point can be constructed from the channel measurements. Our performance guarantees are non-asymptotic, and provide a sufficient condition on the number of samples observed per channel in order to guarantee channel estimates of a certain accuracy. Our analysis uses a model with a "generic" subspace that is drawn at random, and we show the performance bounds hold with high probability. Mathematically, the key estimates are derived by quantifying how well the eigenvectors of certain random matrices approximate the eigenvectors of their mean. We also present a series of numerical results demonstrating that the empirical performance is consistent with the presented theory.

math.NA