SearcharxivSearch

arXiv subjects

Yun Xu

Publications and source records attributed to Yun Xu.

17 recordsLinked to original sources

Efficient Matrix Implementation for Rotary Position Embedding

Rotary Position Embedding (RoPE) has become a core component of modern Transformer architectures across language, vision, and 3D domains. However, existing implementations rely on vector-level split and merge operations that introduce non-negligible computational overhead, often overlooked in attention optimization. The problem is further amplified in multi-dimensional settings (e.g., 2D and 3D RoPE), where additional vector operations and uneven feature partitions degrade hardware utilization. To overcome these limitations, we propose RoME (Rotary Matrix position Embedding), a mathematically equivalent yet computationally efficient reformulation of RoPE that replaces vector operations with unified matrix transformations. RoME eliminates dimension-specific operations, simplifies implementation, and enables fused parallel execution across Cube and Vector units on modern NPUs. Experiments show that RoME delivers substantial acceleration at both the operator and full-model levels. The implementation is available at https://gitcode.com/cann/ops-transformer/blob/master/experimental/posembedding/rope_matrix/README.md.

cs.LG

HiFloat4 Format for Language Model Inference

This paper introduces HiFloat4 (HiF4), a block floating-point data format tailored for deep learning. Each HiF4 unit packs 64 4-bit elements with 32 bits of shared scaling metadata, averaging 4.5 bits per value. The metadata specifies a three-level scaling hierarchy, capturing inter- and intra-group dynamic range while improving the utilization of the representational space. In addition, the large 64-element group size enables matrix multiplications to be executed in a highly fixed-point manner, significantly reducing hardware area and power consumption. To evaluate the proposed format, we conducted inference experiments on several language models, including LLaMA, Qwen, Mistral, DeepSeek-V3.1 and LongCat. Results show that HiF4 achieves higher average accuracy than the state-of-the-art NVFP4 format across multiple models and diverse downstream tasks.

cs.LG

Taiji: A DPU Memory Elasticity Solution for In-production Cloud Environments

The growth of cloud computing drives data centers toward higher density and efficiency. Data processing units (DPUs) enhance server network and storage performance but face challenges such as long hardware upgrade cycles and limited resources. To address these, we propose Taiji, a resource-elasticity architecture for DPUs. Combining hybrid virtualization with parallel memory swapping, Taiji switches the DPU's operating system (OS) into a guest OS and inserts a lightweight virtualization layer, making nearly all DPU memory swappable. It achieves memory overcommitment for the switched guest OS via high-performance memory elasticity, fully transparent to upper-layer applications, and supports hot-switch and hot-upgrade to meet in-production cloud requirements. Experiments show that Taiji expands DPU memory resources by over 50%, maintains virtualization overhead around 5%, and ensures 90% of swap-ins complete within 10 microseconds. Taiji delivers an efficient, reliable, low-overhead elasticity solution for DPUs and is deployed in large-scale production systems across more than 30,000 servers.

cs.OS

Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment

Traditional memory management suffers from metadata overhead, architectural complexity, and stability degradation, problems intensified in cloud environments. Existing software/hardware optimizations are insufficient for cloud computing's dual demands of flexibility and low overhead. This paper presents Vmem, a memory management architecture for in-production cloud environments that enables flexible, efficient cloud server memory utilization through lightweight reserved memory management. Vmem is the first such architecture to support online upgrades, meeting cloud requirements for high stability and rapid iterative evolution. Experiments show Vmem increases sellable memory rate by about 2%, delivers extreme elasticity and performance, achieves over 3x faster boot time for VFIO-based virtual machines (VMs), and improves network performance by about 10% for DPU-accelerated VMs. Vmem has been deployed at large scale for seven years, demonstrating efficiency and stability on over 300,000 cloud servers supporting hundreds of millions of VMs.

cs.OS

Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning

Despite the prosperity of the video language model, the current pursuit of comprehensive video reasoning is thwarted by the inherent spatio-temporal incompleteness within individual videos, resulting in hallucinations and inaccuracies. A promising solution is to augment the reasoning performance with multiple related videos. However, video tokens are numerous and contain redundant information, so directly feeding the relevant video data into a large language model to enhance responses could be counterproductive. To address this challenge, we propose a multi-video collaborative framework for video language models. For efficient and flexible video representation, we establish a Video Structuring Module to represent the video's knowledge as a spatio-temporal graph. Based on the structured video representation, we design the Graph Fusion Module to fuse the structured knowledge and valuable information from related videos into the augmented graph node tokens. Finally, we construct an elaborate multi-video structured prompt to integrate the graph, visual, and textual tokens as the input to the large language model. Extensive experiments substantiate the effectiveness of our framework, showcasing its potential as a promising avenue for advancing video language models. Code will be open-sourced at https://github.com/ziHoHe/SMV-CR.

cs.CV

Analyzing the Impact of Strategic Bidding on the Reserve Capacity via a Bi-Level Model

The growing integration of renewable energy sources necessitates adequate reserve capacity to maintain power balance. However, in market clearing, power companies with flexible resources may submit strategic bids to maximize profits, potentially compromising system reserves. This paper examines the effects of such strategic behavior by modeling the market as a bi-level problem. The upper level represents a strategic company aiming to maximize profit, while the lower level simulates the system operator clearing the market based on submitted offers. To enable duality-based solution methods, we approximate unit commitments with a continuous reserve capacity calculation. Case studies indicate that, in an imperfectly competitive market, more units are incentivized to operate,enhancing system reserves. However, some units go online mainly for profit, ultimately raising electricity costs for consumers. These findings highlight the importance of market design in managing the trade-off between reserve adequacy and economic efficiency in the presence of strategic bidding behavior.

eess.SY

Taylor-Hood like finite elements for nearly incompressible strain gradient elasticity problems

We propose a family of mixed finite elements that are robust for the nearly incompressible strain gradient model, which is a fourth-order singular perturbed elliptic system. The element is similar to [C. Taylor and P. Hood, Comput. & Fluids, 1(1973), 73-100] in the Stokes flow. Using a uniform discrete B-B inequality for the mixed finite element pairs, we show the optimal rate of convergence that is robust in the incompressible limit. By a new regularity result that is uniform in both the materials parameter and the incompressibility, we prove the method converges with $1/2$ order to the solution with strong boundary layer effects. Moreover, we estimate the convergence rate of the numerical solution to the unperturbed second-order elliptic system. Numerical results for both smooth solutions and the solutions with sharp layers confirm the theoretical prediction.

math.NA

Dualities and endomorphisms of pseudo-cones

In this paper we study a class of convex sets which are called closed pseudo-cones and study a new duality of this class. It turns out that the duality characterizes closed pseudo-cones and is essentially the only possible abstract duality of them.The characterization of the duality is corresponding to the classification of endomorphisms closed pseudo-cones.

math.MG

A Derivative-Hilbert operator Acting on Dirichlet spaces

Let $μ$ be a positive Borel measure on the interval $[0,1)$. The Hankel matrix $\mathcal{H}_μ=(μ_{n,k})_{n,k\geq 0}$ with entries $μ_{n,k}=μ_{n+k}$, where $μ_{n}=\int_{[0,1)}t^ndμ(t)$, induces formally the operator as $$\mathcal{DH}_μ(f)(z)=\sum_{n=0}^\infty\left(\sum_{k=0}^\infty μ_{n,k}a_k\right)(n+1)z^n , z\in \mathbb{D},$$ where $f(z)=\sum_{n=0}^{\infty}a_nz^n$ is an analytic function in $\mathbb{D}$. In this paper, we characterize those positive Borel measures on $[0, 1)$ for which $\mathcal{DH}_μ$ is bounded (resp. compact) from Dirichlet spaces $\mathcal{D}_α( 0<α\leq2 )$ into $\mathcal{D}_β( 2\leqβ<4 )$.

math.FA

Challenging Machine Learning-based Clone Detectors via Semantic-preserving Code Transformations

Software clone detection identifies similar code snippets. It has been an active research topic that attracts extensive attention over the last two decades. In recent years, machine learning (ML) based detectors, especially deep learning-based ones, have demonstrated impressive capability on clone detection. It seems that this longstanding problem has already been tamed owing to the advances in ML techniques. In this work, we would like to challenge the robustness of the recent ML-based clone detectors through code semantic-preserving transformations. We first utilize fifteen simple code transformation operators combined with commonly-used heuristics (i.e., Random Search, Genetic Algorithm, and Markov Chain Monte Carlo) to perform equivalent program transformation. Furthermore, we propose a deep reinforcement learning-based sequence generation (DRLSG) strategy to effectively guide the search process of generating clones that could escape from the detection. We then evaluate the ML-based detectors with the pairs of original and generated clones. We realize our method in a framework named CloneGen. CloneGen In evaluation, we challenge the two state-of-the-art ML-based detectors and four traditional detectors with the code clones after semantic-preserving transformations via the aid of CloneGen. Surprisingly, our experiments show that, despite the notable successes achieved by existing clone detectors, the ML models inside these detectors still cannot distinguish numerous clones produced by the code transformations in CloneGen. In addition, adversarial training of ML-based clone detectors using clones generated by CloneGen can improve their robustness and accuracy. CloneGen Meanwhile, compared with the commonly-used heuristics, the DRLSG strategy has shown the best effectiveness in generating code clones to decrease the detection accuracy of the ML-based detectors.

cs.SE

Oxygen Reduction Reaction and X-ray Photoelectron Spectroscopy of Sputtered Fe-N-C Films

Electrocatalysts for the oxygen reduction reaction (ORR) based on complexes of iron and nitrogen in a carbon matrix (Fe-N-C) are a promising alternative to platinum group metal (PGM) based catalysts in polymer electrolyte membrane (PEM) fuel cells. Further improvements of Fe-N-C catalysts would benefit from model thin film studies of activity and stability of catalytic sites, but synthesis of Fe-N-C model thin films is challenging. Here we report on synthesis and characterization of Fe-N-C thin films produced by co-sputtering iron and carbon in a reactive nitrogen atmosphere onto removable glassy carbon rotating disk electrode (RDE) tips. Scanning electron microscopy (SEM) measurements indicate that the Fe-N-C films deposited at high temperature are smoother than the films annealed at high temperature. ORR activity measured on the thin Fe-N-C films is greater for both high-temperature samples than for the room-temperature sample. From the analysis of X-ray photoelectron spectroscopy (XPS) data, exposure of the films to high temperatures results in increased graphitization of the carbon with the Fe-N-C films, and increased relative amount of graphitic and hydrogenated nitrogen species. Overall the results of this study demonstrate the feasibility of a thin film model system approach for studying active sites in PGM-free catalysts.

cond-mat.mtrl-sci

LVMapper: A Large-variance Clone Detector Using Sequencing Alignment Approach

To detect large-variance code clones (i.e. clones with relatively more differences) in large-scale code repositories is difficult because most current tools can only detect almost identical or very similar clones. It will make promotion and changes to some software applications such as bug detection, code completion, software analysis, etc. Recently, CCAligner made an attempt to detect clones with relatively concentrated modifications called large-gap clones. Our contribution is to develop a novel and effective detection approach of large-variance clones to more general cases for not only the concentrated code modifications but also the scattered code modifications. A detector named LVMapper is proposed, borrowing and changing the approach of sequencing alignment in bioinformatics which can find two similar sequences with more differences. The ability of LVMapper was tested on both self-synthetic datasets and real cases, and the results show substantial improvement in detecting large-variance clones compared with other state-of-the-art tools including CCAligner. Furthermore, our new tool also presents good recall and precision for general Type-1, Type-2 and Type-3 clones on the widely used benchmarking dataset, BigCloneBench.

cs.SE

Experimental demonstration of valley-Hall topological photonic crystal at telecommunication wavelengths

Photonic topological insulators provide unprecedented possibilities to eliminate scattering losses and improve the efficiency of optical communication systems. Despite significant theoretical efforts, the experimental demonstration of an integrated photonic topological insulator operating in the telecommunication regime was still missing. Here, we design, fabricate and characterize a photonic-crystal-based topological structure that exhibits valley-Hall effect. We experimentally demonstrate the propagation of topologically protected edge states in a CMOS-compatible chip operating at telecommunication wavelengths. This contribution is an important step towards integrated topological photonics.

physics.optics

Reconfiguring structured light beams using nonlinear metasurfaces

Ultra-compact, low-loss, fast, and reconfigurable optical components, enabling manipulation of light by light, could open numerous opportunities for controlling light on the nanoscale. Nanostructured all-dielectric metasurfaces have been shown to enable extensive control of amplitude and phase of light in the linear optical regime. Among other functionalities, they offer unique opportunities for shaping the wave front of light to introduce the orbital angular momentum (OAM) to a beam. Such structured light beams bring a new degree of freedom for applications ranging from spectroscopy and micromanipulation to classical and quantum optical communications. To date, reconfigurability or tuning of the optical properties of all-dielectric metasurfaces have been achieved mechanically, thermally, electrically or optically, using phase-change or nonlinear optical materials. However, a majority of demonstrated tuning approaches are either slow or require high optical powers. Arsenic trisulfide (As$_2$S$_3$) chalcogenide glass offering ultra-fast and large $χ^{(3)}$ nonlinearity as well as a low two-photon absorption coefficient in the near and mid-wave infrared spectral range, could provide a new platform for the realization of fast and relatively low-intensity reconfigurable metasurfaces. Here, we design and experimentally demonstrate an As$_2$S$_3$ chalcogenide glass based metasurface that enables reshaping of a conventional Hermite-Gaussian beam with no OAM into an OAM beam at low-intensity levels, while preserves the original beam's amplitude and phase characteristics at high-intensity levels. The proposed metasurface could find applications for a new generation of optical communication systems and optical signal processing.

physics.optics

FMtree: A fast locating algorithm of FM-indexes for genomic data

Motivation: As a fundamental task in bioinformatics, searching for massive short patterns over a long text is widely accelerated by various compressed full-text indexes. These indexes are able to provide similar searching functionalities to classical indexes, e.g., suffix trees and suffix arrays, while requiring less space. For genomic data, a well-known family of compressed full-text index, called FM-indexes, presents unmatched performance in practice. One major drawback of FM-indexes is that their locating operations, which report all occurrence positions of patterns in a given text, are particularly slow, especially for the patterns with many occurrences. Results: In this paper, we introduce a novel locating algorithm, FMtree, to fast retrieve all occurrence positions of any pattern via FM-indexes. When searching for a pattern over a given text, FMtree organizes the search space of the locating operation into a conceptual quadtree. As a result, multiple occurrence positions of this pattern can be retrieved simultaneously by traversing the quadtree. Compared with the existing locating algorithms, our tree-based algorithm reduces large numbers of redundant operations and presents better data locality. Experimental results show that FMtree is usually one order of magnitude faster than the state-of-the-art algorithms, and still memory-efficient.

cs.DS

Common Mistakes in Writing Astronomy and Physics Literature in English

This is the 3rd version with major updates and revisions to the 2nd version of the same title. In this article we have collected and corrected some common mistakes made by Chinese students in writing astronomy and physics literature in English. Brief explanations of these mistakes are given in Chinese. We plan to continue to update this collection periodically. Comments, suggestions and criticisms are welcome.

astro-ph.IM

Approaching Gaussian Relay Network Capacity in the High SNR Regime: End-to-End Lattice Codes

We present a natural and low-complexity technique for achieving the capacity of the Gaussian relay network in the high SNR regime. Specifically, we propose the use of end-to-end structured lattice codes with the amplify-and-forward strategy, where the source uses a nested lattice code to encode the messages and the destination decodes the messages by lattice decoding. All intermediate relays simply amplify and forward the received signals over the network to the destination. We show that the end-to-end lattice-coded amplify-and-forward scheme approaches the capacity of the layered Gaussian relay network in the high SNR regime. Next, we extend our scheme to non-layered Gaussian relay networks under the amplify-and-forward scheme, which can be viewed as a Gaussian intersymbol interference (ISI) channel. Compared with other schemes, our approach is significantly simpler and requires only the end-to-end design of the lattice precoding and decoding. It does not require any knowledge of the network topology or the individual channel gains.

cs.IT