SearcharxivSearch

arXiv subjects

Jianhao Xu

Publications and source records attributed to Jianhao Xu.

5 recordsLinked to original sources

Beyond $\ell_2$-norm and $\ell_\infty$-norm: A Curvature-Inspired $\ell_p$-Norm Scheme for Deep Neural Networks

The existing optimizers for deep neural networks (DNNs) typically rely on either the $\ell_2$ norm or the $\ell_\infty$ norm, resulting in optimizers that do not adapt well to substantial changes in curvature across parameter dimensions. Generally, the training process of DNNs often exhibits strong curvature anisotropy in the early period, whereas in the later period, the training process of DNNs tends to move toward flatter regions with weaker anisotropy. Particularly, optimizers based on the \(\ell_2\)-norm are usually dominated by high-curvature directions, restricting updates of optimizers along with lower curvature direction and thus leading to a slower convergence rate. While optimizers based on the \(\ell_\infty\)-norm are prone to oscillations in flatter regions, due to the coordinate-wise updates of the same magnitude. To address these two extreme cases generated by $\ell_2$ and $\ell_\infty$ norms, we propose a novel $\ell_p$-norm scheme with a dynamical value of $p$ and incorporate it into stochastic gradient descent (SGD) and SGD with momentum (SGDM), leading to two novel optimizers with better generalization performance: ${\ell_p}$-SGD (LPSGD) and ${\ell_p}$-SGDM (LPSGDM). Particularly, the resulting optimizers suppress the dominance of high-curvature directions in the early period by utilizing a large $p$ ($p>2$), followed by a gradual decrease of $p$ toward 2 to enable more stable and refined updates, where the latter process is motivated by the cosine annealing strategy. We establish theoretical guarantees of the resulting algorithms and analyze that both LPSGD and LPSGDM achieve an \(O(T^{-1/2})\) convergence rate for the nonconvex setting. Extensive experiments are conducted on benchmark datasets, including CIFAR-10, CIFAR-100, and ImageNet-1K, with multiple DNNs such as VGG-11, ResNet-18, and ResNet-50.

cs.LG

LongCat-Flash-Thinking-2601 Technical Report

We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among open-source models on a wide range of agentic benchmarks, including agentic search, agentic tool use, and tool-integrated reasoning. Beyond benchmark performance, the model demonstrates strong generalization to complex tool interactions and robust behavior under noisy real-world environments. Its advanced capability stems from a unified training framework that combines domain-parallel expert training with subsequent fusion, together with an end-to-end co-design of data construction, environments, algorithms, and infrastructure spanning from pre-training to post-training. In particular, the model's strong generalization capability in complex tool-use are driven by our in-depth exploration of environment scaling and principled task construction. To optimize long-tailed, skewed generation and multi-turn agentic interactions, and to enable stable training across over 10,000 environments spanning more than 20 domains, we systematically extend our asynchronous reinforcement learning framework, DORA, for stable and efficient large-scale multi-environment training. Furthermore, recognizing that real-world tasks are inherently noisy, we conduct a systematic analysis and decomposition of real-world noise patterns, and design targeted training procedures to explicitly incorporate such imperfections into the training process, resulting in improved robustness for real-world applications. To further enhance performance on complex reasoning tasks, we introduce a Heavy Thinking mode that enables effective test-time scaling by jointly expanding reasoning depth and width through intensive parallel thinking.

cs.AI

Machine learning-based prediction of magnet errors in storage ring light sources

Magnet errors in storage rings significantly degrade beam performance, impacting the brightness and stability of the light source. Therefore, beam-based correction is crucial for the safe operation of machines and the stability of radiated photons. Unlike traditional correction methods such as linear optics from closed orbit, this paper proposes a machine learning (ML) framework to directly predict quadrupole/sextupole gradient errors and misalignment from beam position monitor-measured optics functions and closed-orbit distortion data. Based on a four-bend achromat storage ring lattice, we generate training datasets through ELEGANT numerical simulations and compare regression performance of Linear Regression, Support Vector Machine, Radial Basis Function Neural Network and Densely Connected Convolutional Network. Results demonstrate that ML models can effectively predict magnet errors and reconstruct ideal optics. This approach offers a novel strategy for accelerating storage ring commissioning and optimization, online diagnostics, and dynamic compensation for next-generation diffraction-limited rings.

physics.acc-ph

Introducing LongCat-Flash-Thinking: A Technical Report

We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a meticulously crafted training process, beginning with long Chain-of-Thought (CoT) data cold-start and culminating in large-scale Reinforcement Learning (RL). We first employ a well-designed cold-start training strategy, which significantly enhances the reasoning potential and equips the model with specialized skills in both formal and agentic reasoning. Then, a core innovation is our domain-parallel training scheme, which decouples optimization across distinct domains (e.g., STEM, Code, Agentic) and subsequently fuses the resulting expert models into a single, nearly Pareto-optimal model. This entire process is powered by our Dynamic ORchestration for Asynchronous rollout (DORA) system, a large-scale RL framework that delivers a greater than threefold training speedup over synchronous methods on tens of thousands of accelerators. As a result, LongCat-Flash-Thinking achieves state-of-the-art performance among open-source models on a suite of complex reasoning tasks. The model exhibits exceptional efficiency in agentic reasoning, reducing average token consumption by 64.5% (from 19, 653 to 6, 965) on AIME-25, without degrading task accuracy. We release LongCat-Flash-Thinking to promote further advances in reasoning systems and agentic AI research.

cs.AI

Hybrid multi-bend achromat lattice with sextupole cancellation across straight section

The hybrid multi-bend achromat (HMBA) lattice concept is adopted in some diffraction-limited storage ring designs, which can permit relatively large on-momentum dynamic aperture and relatively weak sextupoles. In a typical HMBA lattice, the main arc section is constrained by the transverse phase advances making -I transformation for sextupole cancellation. In this paper, a new HMBA lattice concept with sextupole cancellation across straight section is proposed, where -I is made between adjacent dispersion bumps of two lattice cells. This makes the main arc section free of the phase advance constraint, and as a result, the number of bending magnets (bends) in the lattice cell and the cell tunes can be easily changed, thus providing more choices for lattice design. To achieve the large phase advances required for -I in this new concept, split bend is used as the matching bend, which is a bend split into two pieces with a quadrupole in between. The split bend also serves to reduce the emittance, and the large phase advances also give low beta functions in the straight section enhancing the insertion device brightness. Besides, for a given emittance goal, this new HMBA lattice can have less bends than the typical HMBA lattice due to stronger focusing in bend unit cells, which is beneficial for saving space and suppressing intra-beam scattering effect. Two lattices are given as examples to demonstrate this new concept and show its linear and nonlinear properties, and further extension is also discussed.

physics.acc-ph