SearcharxivSearch

arXiv subjects

Z. F. Wu

Publications and source records attributed to Z. F. Wu.

5 recordsLinked to original sources

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

cs.CL

DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition

We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSeek-V3 to decompose complex problems into a series of subgoals. The proofs of resolved subgoals are synthesized into a chain-of-thought process, combined with DeepSeek-V3's step-by-step reasoning, to create an initial cold start for reinforcement learning. This process enables us to integrate both informal and formal mathematical reasoning into a unified model. The resulting model, DeepSeek-Prover-V2-671B, achieves state-of-the-art performance in neural theorem proving, reaching 88.9% pass ratio on the MiniF2F-test and solving 49 out of 658 problems from PutnamBench. In addition to standard benchmarks, we introduce ProverBench, a collection of 325 formalized problems, to enrich our evaluation, including 15 selected problems from the recent AIME competitions (years 24-25). Further evaluation on these 15 AIME problems shows that the model successfully solves 6 of them. In comparison, DeepSeek-V3 solves 8 of these problems using majority voting, highlighting that the gap between formal and informal mathematical reasoning in large language models is substantially narrowing.

cs.CL

DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

We introduce DeepSeek-Prover-V1.5, an open-source language model designed for theorem proving in Lean 4, which enhances DeepSeek-Prover-V1 by optimizing both training and inference processes. Pre-trained on DeepSeekMath-Base with specialization in formal mathematical languages, the model undergoes supervised fine-tuning using an enhanced formal theorem proving dataset derived from DeepSeek-Prover-V1. Further refinement is achieved through reinforcement learning from proof assistant feedback (RLPAF). Beyond the single-pass whole-proof generation approach of DeepSeek-Prover-V1, we propose RMaxTS, a variant of Monte-Carlo tree search that employs an intrinsic-reward-driven exploration strategy to generate diverse proof paths. DeepSeek-Prover-V1.5 demonstrates significant improvements over DeepSeek-Prover-V1, achieving new state-of-the-art results on the test set of the high school level miniF2F benchmark ($63.5\%$) and the undergraduate level ProofNet benchmark ($25.3\%$).

cs.CL

Proton and molecular permeation through the basal plane of monolayer graphene oxide

Two-dimensional (2D) materials offer a prospect of membranes that combine negligible gas permeability with high proton conductivity and could outperform the existing proton exchange membranes used in various applications including fuel cells. Graphene oxide (GO), a well-known 2D material, facilitates rapid proton transport along its basal plane but proton conductivity across it remains unknown. It is also often presumed that individual GO monolayers contain a large density of nanoscale pinholes that lead to considerable gas leakage across the GO basal plane. Here we show that relatively large, micrometer-scale areas of monolayer GO are impermeable to gases, including helium, while exhibiting proton conductivity through the basal plane which is nearly two orders of magnitude higher than that of graphene. These findings provide insights into the key properties of GO and demonstrate that chemical functionalization of 2D crystals can be utilized to enhance their proton transparency without compromising gas impermeability.

cond-mat.mtrl-sci

Unexpected catalytic activity of nanorippled graphene

Graphite is one of the most chemically inert materials. Its elementary constituent, monolayer graphene, is generally expected to inherit most of the parent material's properties including chemical inertness. Here we show that, unlike graphite, defect-free monolayer graphene exhibits a strong activity with respect to splitting molecular hydrogen, which is comparable to that of metallic and other known catalysts for this reaction. We attribute the unexpected catalytic activity to surface corrugations (nanoscale ripples), a conclusion supported by theory. Nanoripples are likely to play a role in other chemical reactions involving graphene and, because nanorippling is inherent to atomically thin crystals, can be important for two dimensional materials in general.

cond-mat.mtrl-sci