SearcharxivSearch

arXiv subjects

Hexin Wang

Publications and source records attributed to Hexin Wang.

7 recordsLinked to original sources

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers grounded reward signals. The success of code agents illustrates why this matters. As code is executable, compilers and runtimes can provide high-quality rewards for Reinforcement Learning (RL) post-training of LLMs. By contrast, spatial generation still relies largely on fuzzy proxies such as CLIP scores. These signals are fuzzy and biased, making them hard to support RL post-training. Compared with these, game development provides a missing reward environment for spatial world models. A scene encoded by a game engine is an executable world specification: the engine can efficiently check collision, physics, navigability and bounded playability, while the developer provides the global verification signal by judging whether the scene should be accepted. Game development also provides real-world long-horizon trajectory data for RL post-training. We therefore propose Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process.

cs.AI

Stochastic twinning in confined volumes of Mg: Insights from in-situ micromechanical testing and atomistic simulations

Tensile twinning plays a central role in accommodating -axis plasticity in Mg. In bulk Mg, twinning typically shows a relatively deterministic response with a low critical stress, whereas in confined volumes it exhibits pronounced scatter, complicating the prediction of small-scale mechanical behavior. In this study, we investigate the origin of this stochasticity by combining site-specific micropillar compression with atomistic simulations. Experiments show that under -axis compression, plastic deformation is dominated by {10-12} twinning, with each discrete stress drop in the stress-strain response marking the activation and rapid advance of a twin. Atomistic simulations further separate twinning into two mechanistic regimes: nucleation and longitudinal propagation occur in a high-stress, shuffle-assisted regime, whereas lateral thickening proceeds in a low-stress regime controlled by disconnection glide. Linking these mechanistic insights with post-mortem characterization of deformed pillars demonstrates that the scatter in measured yield stresses arises from stochastic selection among competing twinning pathways, governed by the local defect landscape (presence, distribution, and morphology of pre-existing defects). Overall, this work identifies an atomistic basis for size-dependent stochastic twinning in Mg and provides a general framework for materials whose plasticity is controlled by discrete activation events.

cond-mat.mtrl-sci

Pretraining Large Language Models with NVFP4

Large Language Models (LLMs) today are powerful problem solvers across many domains, and they continue to get stronger as they scale in model size, training set size, and training set quality, as shown by extensive research and experimentation across the industry. Training a frontier model today requires on the order of tens to hundreds of yottaflops, which is a massive investment of time, compute, and energy. Improving pretraining efficiency is therefore essential to enable the next generation of even more capable LLMs. While 8-bit floating point (FP8) training is now widely adopted, transitioning to even narrower precision, such as 4-bit floating point (FP4), could unlock additional improvements in computational speed and resource utilization. However, quantization at this level poses challenges to training stability, convergence, and implementation, notably for large-scale models trained on long token horizons. In this study, we introduce a novel approach for stable and accurate training of large language models (LLMs) using the NVFP4 format. Our method integrates Random Hadamard transforms (RHT) to bound block-level outliers, employs a two-dimensional quantization scheme for consistent representations across both the forward and backward passes, utilizes stochastic rounding for unbiased gradient estimation, and incorporates selective high-precision layers. We validate our approach by training a 12-billion-parameter model on 10 trillion tokens -- the longest publicly documented training run in 4-bit precision to date. Our results show that the model trained with our NVFP4-based pretraining technique achieves training loss and downstream task accuracies comparable to an FP8 baseline. These findings highlight that NVFP4, when combined with our training approach, represents a major step forward in narrow-precision LLM training algorithms.

cs.CL

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

We introduce Nemotron-Nano-9B-v2, a hybrid Mamba-Transformer language model designed to increase throughput for reasoning workloads while achieving state-of-the-art accuracy compared to similarly-sized models. Nemotron-Nano-9B-v2 builds on the Nemotron-H architecture, in which the majority of the self-attention layers in the common Transformer architecture are replaced with Mamba-2 layers, to achieve improved inference speed when generating the long thinking traces needed for reasoning. We create Nemotron-Nano-9B-v2 by first pre-training a 12-billion-parameter model (Nemotron-Nano-12B-v2-Base) on 20 trillion tokens using an FP8 training recipe. After aligning Nemotron-Nano-12B-v2-Base, we employ the Minitron strategy to compress and distill the model with the goal of enabling inference on up to 128k tokens on a single NVIDIA A10G GPU (22GiB of memory, bfloat16 precision). Compared to existing similarly-sized models (e.g., Qwen3-8B), we show that Nemotron-Nano-9B-v2 achieves on-par or better accuracy on reasoning benchmarks while achieving up to 6x higher inference throughput in reasoning settings like 8k input and 16k output tokens. We are releasing Nemotron-Nano-9B-v2, Nemotron-Nano12B-v2-Base, and Nemotron-Nano-9B-v2-Base checkpoints along with the majority of our pre- and post-training datasets on Hugging Face.

cs.CL

Solute Co-Segregation Mechanisms at Low-Angle Grain Boundaries in Magnesium: A Combined Atomic-Scale Experimental and Modeling Study

Solute segregation at low-angle grain boundaries (LAGBs) critically affects the microstructure and mechanical properties of magnesium (Mg) alloys. In modern alloys containing multiple substitutional elements, understanding solute-solute interactions at microstructural defects becomes essential for alloy design. This study investigates the co-segregation mechanisms of calcium (Ca), zinc (Zn), and aluminum (Al) at a LAGB in a dilute AZX010 Mg alloy by combining atomic-scale experimental and modeling techniques. Three-dimensional atom probe tomography (3D-APT) revealed significant segregation of Ca, Zn, and Al at the LAGB, with Ca forming linear segregation patterns along dislocation arrays characteristic of the LAGB. Clustering analysis showed increased Ca-Ca pairs at the boundary, indicating synergistic solute interactions. Atomistic simulations and elastic dipole calculations demonstrated that larger Ca atoms prefer tensile regions around dislocations, while smaller Zn and Al atoms favor compressive areas. These simulations also found that Ca-Ca co-segregation near dislocation cores is energetically more favorable than other solute pairings, explaining the enhanced Ca clustering observed experimentally. Thermodynamic modeling incorporating calculated segregation energies and solute-solute interactions accurately predicted solute concentrations at the LAGB, aligning with experimental data. The findings emphasize the importance of solute interactions at dislocation cores in Mg alloys, offering insights for improving mechanical performance through targeted alloying and grain boundary engineering.

cond-mat.mtrl-sci

Grain boundary segregation spectrum in basal-textured Mg alloys: From solute decoration to structural transition

Mg alloys are promising lightweight structural materials due to their low density and excellent mechanical properties. However, their limited formability and ductility necessitate improvements in these properties, specifically through texture modification via grain boundary segregation. While significant efforts have been made, the segregation behavior in Mg polycrystals, particularly with basal texture, remains largely unexplored. In this study, we performed atomistic simulations to investigate grain boundary segregation in dilute and concentrated solid solution Mg-Al alloys. We computed the segregation energy spectrum of basal-textured Mg polycrystals, highlighting the contribution from specific grain boundary sites, such as junctions, and identified a newly discovered bimodal distribution which is distinct compared to the conventional skew-normal distribution found in randomly-oriented polycrystals. Using a hybrid molecular dynamics/Monte Carlo approach, we simulated segregation behavior at finite temperatures, identifying grain boundary structural transitions, particularly the varied fraction and morphology of topologically close-packed grain boundary phases when changing thermodynamic variables. The outcomes of this study offer crucial insights into basal-textured grain boundary segregation and phase formation, which can be extended to other relevant Mg alloys containing topologically close-packed intermetallics.

cond-mat.mtrl-sci

Preliminary Design of CSNS-II Linac SRF LLRF

China Spallation Neutron Source(CSNS) target power will upgrade to 500 kW(CSNS-II) from 300kW, energy gain of H-Linac will up to 300 MeV from 80 MeV using about 50 superconductor cavities. LLRF is an important device for controlling the amplitude and phase of the SRF cavity field to be less than 0.6% and 0.6 deg. The parameters and requirements for CSNS-II Linac LLRF are presented here. The preliminary design work and algorithm verification progress and results at C-ADS Injector-I are introduced.

physics.acc-ph