SearcharxivSearch

arXiv subjects

Xiaodong Yi

Publications and source records attributed to Xiaodong Yi.

16 recordsLinked to original sources

HyQuant: Hybrid-Precision Quantization for LLM Attention

Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low bit-widths, causing performance degradation. Existing methods mainly rely on smoothing techniques to handle outliers, while we propose a hybrid quantization design to better balance accuracy and efficiency. Specifically, we propose \textbf{HyQuant}, an efficient hybrid quantization framework for LLM attention. HyQuant quantizes most attention states into low-bit formats while retaining a small set of vertical-line tokens and local-window states in high precision. These accuracy-critical regions are selected using lightweight vertical-line-aware attention-pattern signals, reducing quantization error with limited overhead. In the Prefill stage, HyQuant uses a hybrid-precision quantized attention operator that preserves vertical-line tokens and a local sliding window in full precision while quantizing the remaining context. In the Decode stage, HyQuant applies the same principle to KV-cache compression and fuses KV dequantization with attention computation to improve memory and hardware efficiency. Across diverse tasks, models, and datasets, HyQuant maintains nearly lossless accuracy with an extremely simple design, demonstrating the efficiency and practical feasibility of hybrid quantization for LLM attention. Code is available at: https://github.com/jerrysfls/HyQuant .

cs.AI

On iterated universal extensions and Nori's fundamental group of nilpotent bundles

Let $k$ be a field of characteristic $0$, $X$ be a geometrically connected, smooth and proper variety over $k$ and $x\in X(k)$ be a base point. Using the notion of iterated universal extensions, we show that Nori's fundamental group $π_{1}^{N}(X,x)$ of nilpotent bundles is uniquely determined by the coherent cohomology groups $\mathrm{H}^{i}(X)=\mathrm{H}^{i}(X,\mathcal{O}_{X})$, $i=1,2$, and the cup product $\cup: \mathrm{H}^{1}(X)\otimes\mathrm{H}^{1}(X) \rightarrow \mathrm{H}^{2}(X)$. This can be seen as an analogue of a classical fact on the de Rham fundamental group of compact Kähler manifolds. As a byproduct, we also determine low degree group cohomology of the trivial representation of $π_{1}^{N}(X,x)$, notably in degree $2$.

math.AG

A finiteness result on representations of Nori's fundamental group scheme

Let $(X,x)$ be a pointed geometrically connected smooth projective variety over a sub-$p$-adic field $K$. For any given rank $n$, we prove that there are only finitely many isomorphism classes of representations $π_{1}^{EF}(X,x)\rightarrow \mathrm{GL}_{n}$, where $π_{1}^{EF}(X,x)$ is Nori's fundamental group of essentially finite bundles. Equivalently, there are only finitely many isomorphism classes of essentially finite bundles of rank $n$. This answers a question from C.Gasbarri.

math.AG

Generating systems, generalized Thomsen collections and derived categories of toric varieties

Bondal claims that for a smooth toric variety $X$, its bounded derived category of coherent sheaves $D_{c}^{b}(X)$ is generated by the Thomsen collection $T(X)$ of line bundles obtained as direct summands of the pushforward of $\mathcal{O}_{X}$ along a Frobenius map with sufficiently divisible degree. The claim is confirmed recently. In this article, we consider a generalized Thomsen collection of line bundles $T(X,D)$ with a $\mathbb{Q}$-divisor $D$ as an auxiliary input, which recovers Thomsen's oringinal collection by setting $D=0$. We introduce the notion of a generating system and prove a theorem on the generation of $\mathcal{O}_{X}$ using many line bundles arising from the generating system. As an application, we verify Bondal's claim for some toric varieties, using a different argument from existing works.

math.AG

A note on cohomological boundedness for $F$-divided sheaves and $\mathcal{D}$-modules

Let $X$ be a smooth proper scheme over an algebraically closed field $k$ in characteristic $p$. In this short note, by interpreting $\mathcal{D}_{X}$-modules as $F$-divided sheaves and establishing a cohomological boundedness property for $F$-divided sheaves, we prove that any $\mathcal{O}_{X}$-coherent $\mathcal{D}_{X}$-module has finite dimensional $\mathcal{D}_{X}$-module cohomology.

math.AG

On the Tannakian group scheme of $\mathcal{D}$-modules

We prove a Lefschetz theorem for the Tannakian group scheme of $\mathcal{D}$-modules, in arbitrary characteristic. In characteristic $0$, We prove a Künneth formula for the Tannakian group scheme of regular singular integrable connections, and disprove it for the Tannakian group scheme of all integrable connections without any regularity assumption.

math.AG

Task2Morph: Differentiable Task-inspired Framework for Contact-Aware Robot Design

Optimizing the morphologies and the controllers that adapt to various tasks is a critical issue in the field of robot design, aka. embodied intelligence. Previous works typically model it as a joint optimization problem and use search-based methods to find the optimal solution in the morphology space. However, they ignore the implicit knowledge of task-to-morphology mapping which can directly inspire robot design. For example, flipping heavier boxes tends to require more muscular robot arms. This paper proposes a novel and general differentiable task-inspired framework for contact-aware robot design called Task2Morph. We abstract task features highly related to task performance and use them to build a task-to-morphology mapping. Further, we embed the mapping into a differentiable robot design process, where the gradient information is leveraged for both the mapping learning and the whole optimization. The experiments are conducted on three scenarios, and the results validate that Task2Morph outperforms DiffHand, which lacks a task-inspired morphology module, in terms of efficiency and effectiveness.

cs.RO

A Study on Training and Developing Large Language Models for Behavior Tree Generation

This paper presents an innovative exploration of the application potential of large language models (LLM) in addressing the challenging task of automatically generating behavior trees (BTs) for complex tasks. The conventional manual BT generation method is inefficient and heavily reliant on domain expertise. On the other hand, existing automatic BT generation technologies encounter bottlenecks related to task complexity, model adaptability, and reliability. In order to overcome these challenges, we propose a novel methodology that leverages the robust representation and reasoning abilities of LLMs. The core contribution of this paper lies in the design of a BT generation framework based on LLM, which encompasses the entire process, from data synthesis and model training to application developing and data verification. Synthetic data is introduced to train the BT generation model (BTGen model), enhancing its understanding and adaptability to various complex tasks, thereby significantly improving its overall performance. In order to ensure the effectiveness and executability of the generated BTs, we emphasize the importance of data verification and introduce a multilevel verification strategy. Additionally, we explore a range of agent design and development schemes with LLM as the central element. We hope that the work in this paper may provide a reference for the researchers who are interested in BT generation based on LLMs.

cs.CL

MVP: Meta Visual Prompt Tuning for Few-Shot Remote Sensing Image Scene Classification

Vision Transformer (ViT) models have recently emerged as powerful and versatile models for various visual tasks. Recently, a work called PMF has achieved promising results in few-shot image classification by utilizing pre-trained vision transformer models. However, PMF employs full fine-tuning for learning the downstream tasks, leading to significant overfitting and storage issues, especially in the remote sensing domain. In order to tackle these issues, we turn to the recently proposed parameter-efficient tuning methods, such as VPT, which updates only the newly added prompt parameters while keeping the pre-trained backbone frozen. Inspired by VPT, we propose the Meta Visual Prompt Tuning (MVP) method. Specifically, we integrate the VPT method into the meta-learning framework and tailor it to the remote sensing domain, resulting in an efficient framework for Few-Shot Remote Sensing Scene Classification (FS-RSSC). Furthermore, we introduce a novel data augmentation strategy based on patch embedding recombination to enhance the representation and diversity of scenes for classification purposes. Experiment results on the FS-RSSC benchmark demonstrate the superior performance of the proposed MVP over existing methods in various settings, such as various-way-various-shot, various-way-one-shot, and cross-domain adaptation.

cs.CV

Expediting Distributed DNN Training with Device Topology-Aware Graph Deployment

This paper presents TAG, an automatic system to derive optimized DNN training graph and its deployment onto any device topology, for expedited training in device- and topology- heterogeneous ML clusters. We novelly combine both the DNN computation graph and the device topology graph as input to a graph neural network (GNN), and join the GNN with a search-based method to quickly identify optimized distributed training strategies. To reduce communication in a heterogeneous cluster, we further explore a lossless gradient compression technique and solve a combinatorial optimization problem to automatically apply the technique for training time minimization. We evaluate TAG with various representative DNN models and device topologies, showing that it can achieve up to 4.56x training speed-up as compared to existing schemes. TAG can produce efficient deployment strategies for both unseen DNN models and unseen device topologies, without heavy fine-tuning.

cs.LG

Optimizing DNN Compilation for Distributed Training with Joint OP and Tensor Fusion

This paper proposes DisCo, an automatic deep learning compilation module for data-parallel distributed training. Unlike most deep learning compilers that focus on training or inference on a single device, DisCo optimizes a DNN model for distributed training over multiple GPU machines. Existing single-device compilation strategies do not work well in distributed training, due mainly to communication inefficiency that they incur. DisCo generates optimized, joint computation operator and communication tensor fusion strategies to enable highly efficient distributed training. A GNN-based simulator is built to effectively estimate per-iteration training time achieved by operator/tensor fusion candidates. A backtracking search algorithm is driven by the simulator, navigating efficiently in the large strategy space to identify good operator/tensor fusion strategies that minimize distributed training time. We compare DisCo with existing DL fusion schemes and show that it achieves good training speed-up close to the ideal, full computation-communication overlap case.

cs.DC

Efficient Pipeline Planning for Expedited Distributed DNN Training

To train modern large DNN models, pipeline parallelism has recently emerged, which distributes the model across GPUs and enables different devices to process different microbatches in pipeline. Earlier pipeline designs allow multiple versions of model parameters to co-exist (similar to asynchronous training), and cannot ensure the same model convergence and accuracy performance as without pipelining. Synchronous pipelining has recently been proposed which ensures model performance by enforcing a synchronization barrier between training iterations. Nonetheless, the synchronization barrier requires waiting for gradient aggregation from all microbatches and thus delays the training progress. Optimized pipeline planning is needed to minimize such wait and hence the training time, which has not been well studied in the literature. This paper designs efficient, near-optimal algorithms for expediting synchronous pipeline-parallel training of modern large DNNs over arbitrary inter-GPU connectivity. Our algorithm framework comprises two components: a pipeline partition and device mapping algorithm, and a pipeline scheduler that decides processing order of microbatches over the partitions, which together minimize the per-iteration training time. We conduct thorough theoretical analysis, extensive testbed experiments and trace-driven simulation, and demonstrate our scheme can accelerate training up to 157% compared with state-of-the-art designs.

cs.DC

OneFlow: Redesign the Distributed Deep Learning Framework from Scratch

Deep learning frameworks such as TensorFlow and PyTorch provide a productive interface for expressing and training a deep neural network (DNN) model on a single device or using data parallelism. Still, they may not be flexible or efficient enough in training emerging large models on distributed devices, which require more sophisticated parallelism beyond data parallelism. Plugins or wrappers have been developed to strengthen these frameworks for model or pipeline parallelism, but they complicate the usage and implementation of distributed deep learning. Aiming at a simple, neat redesign of distributed deep learning frameworks for various parallelism paradigms, we present OneFlow, a novel distributed training framework based on an SBP (split, broadcast and partial-value) abstraction and the actor model. SBP enables much easier programming of data parallelism and model parallelism than existing frameworks, and the actor model provides a succinct runtime mechanism to manage the complex dependencies imposed by resource constraints, data movement and computation in distributed deep learning. We demonstrate the general applicability and efficiency of OneFlow for training various large DNN models with case studies and extensive experiments. The results show that OneFlow outperforms many well-known customized libraries built on top of the state-of-the-art frameworks. The code of OneFlow is available at: https://github.com/Oneflow-Inc/oneflow.

cs.DC

NFVactor: A Resilient NFV System using the Distributed Actor Model

Resilience functionality, including failure resilience and flow migration, is of pivotal importance in practical network function virtualization (NFV) systems. However, existing failure recovery procedures incur high packet processing delay due to heavyweight process-checkpointing, while flow migration has poor performance due to centralized control. This paper proposes NFVactor, a novel NFV system that aims to provide lightweight failure resilience and high-performance flow migration. NFVactor enables these by using actor model to provide a per-flow execution environment, so that each flow can replicate and migrate itself with improved parallelism, while the efficiency of the actor model is guaranteed by a carefully designed runtime system. Moreover, NFVactor achieves transparent resilience: once a new NF is implemented for NFVactor, the NF automatically acquires resilience support. Our evaluation result shows that NFVactor achieves 10Gbps packet processing, flow migration completion time that is 197 times faster than existing system, and packet processing delay stablized at around 20 microseconds during replication.

cs.NI

CNN Feature boosted SeqSLAM for Real-Time Loop Closure Detection

Loop closure detection (LCD) is an indispensable part of simultaneous localization and mapping systems (SLAM); it enables robots to produce a consistent map by recognizing previously visited places. When robots operate over extended periods, robustness to viewpoint and condition changes as well as satisfactory real-time performance become essential requirements for a practical LCD system. This paper presents an approach to directly utilize the outputs at the intermediate layer of a pre-trained convolutional neural network (CNN) as image descriptors. The matching location is determined by matching the image sequences through a method called SeqCNNSLAM. The utility of SeqCNNSLAM is comprehensively evaluated in terms of viewpoint and condition invariance. Experiments show that SeqCNNSLAM outperforms state-of-the-art LCD systems, such as SeqSLAM and Change Removal, in most cases. To allow for the real-time performance of SeqCNNSLAM, an acceleration method, A-SeqCNNSLAM, is established. This method exploits the location relationship between the matching images of adjacent images to reduce the matching range of the current image. Results demonstrate that acceleration of 4-6 is achieved with minimal accuracy degradation, and the method's runtime satisfies the real-time demand. To extend the applicability of A-SeqCNNSLAM to new environments, a method called O-SeqCNNSLAM is established for the online adjustment of the parameters of A-SeqCNNSLAM.

cs.RO

Joint Communication-Motion Planning in Wireless-Connected Robotic Networks: Overview and Design Guidelines

Recent years have witnessed the prosperity of robots and in order to support consensus and cooperation for multi-robot system, wireless communications and networking among robots and the infrastructure have become indispensable. In this technical note, we first provide an overview of the research contributions on communication-aware motion planning (CAMP) in designing wireless-connected robotic networks (WCRNs), where the degree-of-freedom (DoF) provided by motion and communication capabilities embraced by the robots have not been fully exploited. Therefore, we propose the framework of joint communication-motion planning (JCMP) as well as the architecture for incorporating JCMP in WCRNs. The proposed architecture is motivated by the observe-orient-decision-action (OODA) model commonly adopted in robotic motion control and cognitive radio. Then, we provide an overview of the orient module that quantify the connectivity assessment. Afterwards, we highlight the JCMP module and compare it with the conventional communication-planning, where the necessity of the JCMP is validated via both theoretical analysis and simulation results of an illustrative example. Finally, a series of open problems are discussed, which picture the gap between the state-of-the-art and a practical WCRN.

cs.RO