SearcharxivSearch

arXiv subjects

Weiqiang Wu

Publications and source records attributed to Weiqiang Wu.

6 recordsLinked to original sources

MoE Proxy Models for Low-Cost Failure Reproduction and Diagnosis in LLM RL Post-Training

Reinforcement learning (RL) post-training of large language models (LLMs) is computationally intensive and involves complex system pipelines with substantial debugging overhead. In practice, factors such as framework adaptation, numerical precision, and operator implementation can cause failures, including gradient overflow and loss divergence. Reproducing such failures directly on large models requires considerable time and computational resources. This paper systematically analyzes failures encountered during large-scale RL training on the Huawei Ascend platform, summarizes representative failure types, and identifies three model-side factors relevant to fault reproduction. Based on these factors, we propose a proxy-model construction method for low-cost fault investigation and auxiliary diagnosis. It employs structure-preserving, clustering-based expert pruning to select representative experts while retaining the model's backbone architecture, routing mechanism, and basic task capabilities. Our experimental results show that the proxy models reduce accelerator requirements by 50%-87.5% and achieve up to a 33.3x reduction in per-step NPU-hour cost, while preserving major training dynamics and reproducing fault responses consistent with the original models. Overall, the proxy models can serve as low-cost surrogates for fault reproduction, targeted validation, and auxiliary diagnosis in RL post-training.

cs.LG

Federated Linear Contextual Bandits

This paper presents a novel federated linear contextual bandits model, where individual clients face different $K$-armed stochastic bandits coupled through common global parameters. By leveraging the geometric structure of the linear rewards, a collaborative algorithm called Fed-PE is proposed to cope with the heterogeneity across clients without exchanging local feature vectors or raw data. Fed-PE relies on a novel multi-client G-optimal design, and achieves near-optimal regrets for both disjoint and shared parameter cases with logarithmic communication costs. In addition, a new concept called collinearly-dependent policies is introduced, based on which a tight minimax regret lower bound for the disjoint parameter case is derived. Experiments demonstrate the effectiveness of the proposed algorithms on both synthetic and real-world datasets.

stat.ML

K-Core based Temporal Graph Convolutional Network for Dynamic Graphs

Graph representation learning is a fundamental task in various applications that strives to learn low-dimensional embeddings for nodes that can preserve graph topology information. However, many existing methods focus on static graphs while ignoring evolving graph patterns. Inspired by the success of graph convolutional networks(GCNs) in static graph embedding, we propose a novel k-core based temporal graph convolutional network, the CTGCN, to learn node representations for dynamic graphs. In contrast to previous dynamic graph embedding methods, CTGCN can preserve both local connective proximity and global structural similarity while simultaneously capturing graph dynamics. In the proposed framework, the traditional graph convolution is generalized into two phases, feature transformation and feature aggregation, which gives the CTGCN more flexibility and enables the CTGCN to learn connective and structural information under the same framework. Experimental results on 7 real-world graphs demonstrate that the CTGCN outperforms existing state-of-the-art graph embedding methods in several tasks, including link prediction and structural role classification. The source code of this work can be obtained from \url{https://github.com/jhljx/CTGCN}.

cs.LG

Stochastic Linear Contextual Bandits with Diverse Contexts

In this paper, we investigate the impact of context diversity on stochastic linear contextual bandits. As opposed to the previous view that contexts lead to more difficult bandit learning, we show that when the contexts are sufficiently diverse, the learner is able to utilize the information obtained during exploitation to shorten the exploration process, thus achieving reduced regret. We design the LinUCB-d algorithm, and propose a novel approach to analyze its regret performance. The main theoretical result is that under the diverse context assumption, the cumulative expected regret of LinUCB-d is bounded by a constant. As a by-product, our results improve the previous understanding of LinUCB and strengthen its performance guarantee.

cs.LG

On Embedded Spheres of Affine Manifolds

This paper studies certain embedded spheres in closed affine manifolds. For $n \geq 3$, we investigate the dome bodies in a closed affine $n$-manifold $M$ with its boundary homeomorphic to a sphere under the assumption that a developing map restricted to a component of $\partial\hat{M}$ is an embedding onto a strictly convex sphere in $\mathbb{A}^n$. By using the recurrent property of an incomplete geodesic we show that dome bodies are compact. Then a maximal dome body is a closed solid ball bounded by a component of $\partial\hat{M}$, and hence equals $\hat{M}$. The main theorem is that the standard ball in an affine space can only bound one compact affine manifold inside, namely the solid ball.

math.GT

On a Closed Binding Curve of One-holed Torus

Given a closed binding curve $γ$ of a surface $Σ$, any equivalence class of marked complete hyperbolic structure can be decomposed into polygons(possibly with a puncture) with sides being hyperbolic geodesic segments. When $Σ$ is a one-holed torus and $γ= A^3 B^2$, we show that any equivalence class of marked complete hyperbolic structure gives rise to an equilateral bigon with a puncture and a hexagon with equal opposite sides. In particular, we give a new coordinates of the Fricke Space of the one-holed torus.

math.GT