SearcharxivSearch

arXiv subjects

Qiao Zhang

Publications and source records attributed to Qiao Zhang.

At least 19 recordsLinked to original sources

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online-softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware-aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped-query attention, variable-length batches, and INT8 key/value caches. Across LongEval, LV-Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long-context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy-efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang-Ulrica/TileMix.

cs.AI

Multiple Vehicles and Traction Network Interaction System Stability Analysis and Oscillation Responsibility Identification

The electrical incompatibility between vehicles and traction network in railway system can result in system instability and oscillation overvoltage issues. To analyze the system stability, impedance-based frequency-domain methods are commonly used. However, the current impedance-based modeling methods face challenges in practical implementation due to the requirement of precise analytical models and detailed internal parameters for all vehicles. Moreover, multiple vehicles operate simultaneously in railway systems, each with different operating conditions and internal parameters, thereby influencing system stability to different extents. Therefore, it is crucial to accurately identify the critical vehicles to prevent resonance accidents. To address these challenges, a component connection-based modeling approach for the railway vehicle-grid system is proposed, which only requires the measured impedance results without the internal information of vehicles. In addition, a multilevel sensitivity analysis method is introduced to quantitatively identify the critical vehicles and internal parameters that influence system stability, which outperforms traditional sensitivity analysis methods in computational complexity. Furthermore, a system-level electrical compatibility test process for the railway vehicle-grid system is provided, incorporating the proposed stability and sensitivity analysis methods. Finally, case studies based on the real-world train schedule of a multivehicle-accessed railway vehicle-grid system are designed to verify the correctness of the proposed method.

eess.SY

A Flow Model for the Electrified Railway-Power Grid Hybrid Asymmetric Coupled System and its Linearized Method

In mountainous regions where traction loads constitute a significant portion of a long-chain weak power grid (PG) with sustainable energy, the interaction between the traction power supply system and the PG becomes increasingly evident. The integrated power flow calculation (PFC) method and its linearized model are quite important for the PG - traction network (TN) joint planning. However, existing research on the port load characteristics of the EMUs and the connection angle characteristics of traction transformers is insufficient, and there is a lack of effective methods for PFC or linearized PFC in systems that couple the PG with the traction network. To fill this gap, this paper proposes an integrated PFC model for the AT TN - PG coupled system, along with a linearized method. Firstly, according to the relationship of the phases between the PG and the AT traction network, the node admittance matrix of the coupled system has been constructed. Then, the issue of power injection equations being unable to deal with the EMUs port load is resolved by merging the contact line node and the rail node. Subsequently, the integrated PFC equations for the coupling system are established. Next, a hybrid phase linear decoupled power flow model for the coupling system is developed, employing the correspondence between the phases of the PG and the TN, as well as the phase angle differences among various nodes and branches. Numerical simulations conducted in a specific region demonstrate the necessity of an integrated PFC for the coupled system and validate both the accuracy and efficiency of the linearized model.

eess.SY

An Improved Deep Reinforcement Learning Control Strategy for Traction Dual Rectifiers in EMUs

Due to the use of PI-based d q current decoupling in the pulse rectifier of CRH5 high-speed trains, the PI parameters directly affect the traction system's control performance. Linearized control may have issues with reference trajectory changes or model mismatches, leading to a decrease in system performance, while nonlinear control may have problems with jitter and poor steady-state accuracy. This paper proposes a new control strategy that replaces all PI in the d q current decoupling control with a single intelligent agent. This method based on Deep Reinforcement Learning (DRL) can avoid various drawbacks of linearization and nonlinear control and ensure the stability of intermediate DC voltage. However, when EMUs are in different working conditions and switching, the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm used in traction dual rectifiers does not have a good control effect. Focusing on the issue, Reward Shaping (RS) is added to re-design a nonlinear reward function, which can be combined with Prioritized Experience Replay (PER) to increase the convergence speed of the episode reward. The simulation results show that the improved control strategy can be effectively applied to EMUs working in multiple conditions. Finally, the stability analysis is carried out using Lyapunov's second method and the verification results of the hardware-in-the-loop (HIL) simulation platform show that the DRL control has a good effect.

eess.SY

Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.

cs.CV

PRAG: End-to-End Privacy-Preserving Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) is essential for enhancing Large Language Models (LLMs) with external knowledge, but its reliance on cloud environments exposes sensitive data to privacy risks. Existing privacy-preserving solutions often sacrifice retrieval quality due to noise injection or only provide partial encryption. We propose PRAG, an end-to-end privacy-preserving RAG system that achieves end-to-end confidentiality for both documents and queries without sacrificing the scalability of cloud-hosted RAG. PRAG features a dual-mode architecture: a non-interactive PRAG-I utilizes homomorphic-friendly approximations for low-latency retrieval, while an interactive PRAG-II leverages client assistance to match the accuracy of non-private RAG. To ensure robust semantic ordering, we introduce Operation-Error Estimation (OEE), a mechanism that stabilizes ranking against homomorphic noise. Experiments on large-scale datasets demonstrate that PRAG achieves competitive recall (72.45%-74.45%), practical retrieval latency, and strong resilience against graph reconstruction attacks while maintaining end-to-end confidentiality. This work confirms the feasibility of secure, high-performance RAG at scale.

cs.CR

Petabit-per-second Random Number Generation

Physical random number generators based on chaotic microcombs, with their complex nonlinear dynamics and multi-channel parallel capability, have attracted considerable research attention. However, key technical challenges for chaotic microcombs are the high correlation between symmetric teeth and the low bandwidth of single-channel teeth, which seriously affect the speed and scalability of random number generation. We experimentally demonstrate a petabit-per-second (Pbit/s) parallel random number generation system based on intensity chaotic modulation and Rayleigh scattering. Through intensity modulation, the effective bandwidth of the single-channel entropy source is increased from 440MHz to 27.6GHz. Crucially, Rayleigh scattering further contributes through the random superposition of backscattered light, which introduces unpredictable fluctuations in intensity, phase, and polarization. This randomness suppresses inter-channel correlation among parallel entropy sources to ~0.02, ensuring their orthogonality. Moreover, by employing polarization-diverse coherent detection on a single-channel, four new low correlated sub-channels are extracted: X-/Y- intensity and phase. We achieve a single-channel bit rate of 14.336 Tbit/s and a total bit rate of 1.032 Pbit/s (over 72 parallel channels) with offline post-processing, representing the highest post-processing record reported in both the single-channel and the total system. Moreover, our scheme based on a single chaotic microcomb and fiber scattering link show fundamentally scalable. The total bit rate can be significantly pushed beyond the Pbit/s level by further expanding the usable comb channel and/or by deploying multiple fiber scattering links in parallel, paving a practical path toward higher throughput regimes.

physics.optics

SecDTD: Dynamic Token Drop for Secure Transformers Inference

The rapid adoption of Transformer-based AI has been driven by accessible models such as ChatGPT, which provide API-based services for developers and businesses. However, as these online inference services increasingly handle sensitive inputs, privacy concerns have emerged as a significant challenge. To address this, secure inference frameworks have been proposed, but their high computational and communication overhead often limit practical deployment. In plaintext settings, token drop is an effective technique for reducing inference cost; however, our analysis reveals that directly applying such methods to ciphertext scenarios is suboptimal due to distinct cost distributions in secure computation. We propose SecDTD, a dynamic token drop scheme tailored for secure Transformer inference. SecDTD advances token drop by shifting the dropping to earlier inference stages, effectively reducing the cost of key components such as Softmax. To support this, we introduce two core techniques. Max-Centric Normalization (MCN): A novel, Softmax-independent scoring method that enables early token drop with minimal overhead and improved normalization, supporting more aggressive dropping without accuracy loss. OMSel: A faster, oblivious median selection protocol that securely identifies the median of importance scores to support token drop. Compared to existing sorting-based methods, OMSel achieves a 16.9$\times$ speedup while maintaining security, obliviousness and randomness. We evaluate SecDTD through 48 experiments across eight GLUE datasets under various network settings using the BOLT and BumbleBee frameworks. SecDTD achieves 4.47 times end-to-end inference acceleration without degradation in accuracy.

cs.CR

Almost-Free Queue Jumping for Prior Inputs in Private Neural Inference

Privacy-Preserving Machine Learning as a Service (PP-MLaaS) enables secure neural network inference by integrating cryptographic primitives such as homomorphic encryption (HE) and multi-party computation (MPC), protecting both client data and server models. Recent mixed-primitive frameworks have significantly improved inference efficiency, yet they process batched inputs sequentially, offering little flexibility for prioritizing urgent requests. Naïve queue jumping introduces considerable computational and communication overhead, increasing non-negligible latency for in-queue inputs. We initiate the study of privacy-preserving queue jumping in batched inference and propose PrivQJ, a novel framework that enables efficient priority handling without degrading overall system performance. PrivQJ exploits shared computation across inputs via in-processing slot recycling, allowing prior inputs to be piggybacked onto ongoing batch computation with almost no additional cryptographic cost. Both theoretical analysis and experimental results demonstrate over an order-of-magnitude reduction in overhead compared to state-of-the-art PP-MLaaS systems.

cs.CR

Denoising diffusion and latent diffusion models for physics field simulations

Accurate prediction of physical fields is critical in various engineering applications, including thermal management in electronic systems, airfoil shape optimization in aerospace, and flow field control in hypersonic vehicles. This study employs the Denoising Diffusion Probabilistic Models (DDPMs) for predicting the temperature field caused by the thermal diffusion, and the flow fields spanning from incompressible to hypersonic regimes. A conditional DDPM framework is first validated with a steady-state thermal diffusion problem by predicting the temperature distribution around a plate with holes. Strong agreement with ground truth data is shown with an average error of approximately 0.013 for plates with a central circular hole. The model also delivers high accuracy in critical regions, such as near the inner circular or square holes. Its performance is further evaluated on incompressible flow around an airfoil and hypersonic flow over a compression ramp, confirming robust predictive capability across diverse flow conditions. Additionally, a latent-space implementation of DDPM is introduced, which employs an Autoencoder (AE) for dimensionality reduction and reconstruction of the physical data. The resulting Latent Diffusion Model (LDM) maintains reconstruction quality comparable to the standard DDPM while substantially reducing the computational cost of the diffusion training process. When applied to hypersonic flow over a compression ramp in the original parameter space, LDM predictions align well with ground truth, achieving a deviation of only 4.28% in separation length estimation. This work confirms the high predictive accuracy of the DDPM framework and highlights the efficiency gains from performing diffusion in a learned latent space. The findings establish an efficient framework for high fidelity generative modeling of complex thermal/flow fields.

physics.flu-dyn

Towards Zero Rotation and Beyond: Architecting Neural Networks for Fast Secure Inference with Homomorphic Encryption

Privacy-preserving deep learning addresses privacy concerns in Machine Learning as a Service (MLaaS) by using Homomorphic Encryption (HE) for linear computations. However, the computational overhead remains a major challenge. While prior work has improved efficiency, most approaches build on models originally designed for plaintext inference. Such models incur architectural inefficiencies when adapted to HE. We argue that substantial gains require networks tailored to HE rather than retrofitting plaintext architectures. Our design has two components: the building block and the overall architecture. First, StriaBlock targets the most expensive HE operation, rotation. It integrates ExRot-Free Convolution and a novel Cross Kernel, eliminating external rotations and requiring only 19% of the internal rotations used by plaintext models. Second, our architectural principles include (i) the Focused Constraint Principle, which limits cost-sensitive factors while preserving flexibility elsewhere, and (ii) the Channel Packing-Aware Scaling Principle, which adapts bottleneck ratios to ciphertext channel capacity that varies with depth. Together, these strategies control both local and end-to-end HE cost, enabling a balanced HE-tailored network. We evaluate the resulting StriaNet across datasets of varying scales, including ImageNet, Tiny ImageNet, and CIFAR-10. At comparable accuracy, StriaNet achieves speedups of 9.78x, 6.01x, and 9.24x on ImageNet, Tiny ImageNet, and CIFAR-10, respectively.

cs.CR

HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation

We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion Transformer (DiT)-based flow matching models to the billion-parameter scale within the motion generation domain, delivering instruction-following capabilities that significantly outperform current open-source benchmarks. Uniquely, we introduce a comprehensive, full-stage training paradigm -- including large-scale pretraining on over 3,000 hours of motion data, high-quality fine-tuning on 400 hours of curated data, and reinforcement learning from both human feedback and reward models -- to ensure precise alignment with the text instruction and high motion quality. This framework is supported by our meticulous data processing pipeline, which performs rigorous motion cleaning and captioning. Consequently, our model achieves the most extensive coverage, spanning over 200 motion categories across 6 major classes. We release HY-Motion 1.0 to the open-source community to foster future research and accelerate the transition of 3D human motion generation models towards commercial maturity.

cs.CV

Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach

General Matrix Multiplication (GEMM) is a critical operation underpinning a wide range of applications in high-performance computing (HPC) and artificial intelligence (AI). The emergence of hardware optimized for low-precision arithmetic necessitates a reevaluation of numerical algorithms to leverage mixed-precision computations, achieving improved performance and energy efficiency. This research introduces an adaptive mixed-precision GEMM framework that supports different precision formats at fine-grained tile/block levels. We utilize the PaRSEC runtime system to balance workloads across various architectures. The performance scales well on ARM CPU-based Fugaku supercomputer, Nvidia GPU-based A100 DGX, and AMD GPU-based Frontier supercomputer. This research aims to enhance computational efficiency and accuracy by bridging algorithmic advancements and hardware innovations, driving transformative progress in various applications.

cs.DC

Maximum Shortest Path Interdiction Problem by Upgrading Nodes on Trees under Unit Cost

Network interdiction problems by deleting critical nodes have wide applications. However, node deletion is not always feasible in certain practical scenarios. We consider the maximum shortest path interdiction problem by upgrading nodes on trees under unit cost (MSPIT-UN$_u$). It aims to upgrade a subset of nodes to maximize the length of the shortest root-leaf distance such that the total upgrade cost under unit cost is upper bounded by a given value. We develop a dynamic programming algorithm with a time complexity of $O(n^3)$ to solve this problem. Furthermore, we consider the related minimum cost problem of (MSPIT-UN$_u$) and propose an $O(n^3\log n)$ binary search algorithm, where a dynamic programming algorithm is exceeded in each iteration to solve its corresponding problem (MSPIT-UN$_u$). Finally, we design numerical experiments to show the effectiveness of the algorithms.

math.OC

On the Expressive Power of Subgraph Graph Neural Networks for Graphs with Bounded Cycles

Graph neural networks (GNNs) have been widely used in graph-related contexts. It is known that the separation power of GNNs is equivalent to that of the Weisfeiler-Lehman (WL) test; hence, GNNs are imperfect at identifying all non-isomorphic graphs, which severely limits their expressive power. This work investigates $k$-hop subgraph GNNs that aggregate information from neighbors with distances up to $k$ and incorporate the subgraph structure. We prove that under appropriate assumptions, the $k$-hop subgraph GNNs can approximate any permutation-invariant/equivariant continuous function over graphs without cycles of length greater than $2k+1$ within any error tolerance. We also provide an extension to $k$-hop GNNs without incorporating the subgraph structure. Our numerical experiments on established benchmarks and novel architectures validate our theory on the relationship between the information aggregation distance and the cycle size.

cs.LG

The Restricted Inverse Optimal Value Problem under Weighted Bottle-neck Hamming distance on trees

We consider the Restricted Inverse Optimal Value Problem (RIOVSP) on trees under weighted bottleneck Hamming distance, denoted as (RIOVSPT$_{BH}$). The problem aims to minimize the total cost under weighted bottle-neck Hamming distance such that the length of the shortest root-leaf path of the tree is lower-bounded by a given value by adjusting the length of some edges. Additionally, the specified lower bound must correspond to the length of a particular root-leaf path. Through careful analysis of the problem's structural properties, we develop an algorithm with $O(n\log n)$ time complexity to solve (RIOVSPT$_{BH}$). Furthermore, by removing the path-length constraint, we derive the Minimum Cost Shortest Path Interdiction Problem on Trees (MCSPIT), for which we present an $O(n\log n)$ time algorithm that operates under weighted bottleneck Hamming distance. Extensive computational experiments demonstrate the efficiency and effectiveness of both algorithms.

cs.DS

FirmRCA: Towards Post-Fuzzing Analysis on ARM Embedded Firmware with Efficient Event-based Fault Localization

While fuzzing has demonstrated its effectiveness in exposing vulnerabilities within embedded firmware, the discovery of crashing test cases is only the first step in improving the security of these critical systems. The subsequent fault localization process, which aims to precisely identify the root causes of observed crashes, is a crucial yet time-consuming post-fuzzing work. Unfortunately, the automated root cause analysis on embedded firmware crashes remains an underexplored area, which is challenging from several perspectives: (1) the fuzzing campaign towards the embedded firmware lacks adequate debugging mechanisms, making it hard to automatically extract essential runtime information for analysis; (2) the inherent raw binary nature of embedded firmware often leads to over-tainted and noisy suspicious instructions, which provides limited guidance for analysts in manually investigating the root cause and remediating the underlying vulnerability. To address these challenges, we design and implement FirmRCA, a practical fault localization framework tailored specifically for embedded firmware. FirmRCA introduces an event-based footprint collection approach to aid and significantly expedite reverse execution. Next, to solve the complicated memory alias problem, FirmRCA proposes a history-driven method by tracking data propagation through the execution trace, enabling precise identification of deep crash origins. Finally, FirmRCA proposes a novel strategy to highlight key instructions related to the root cause, providing practical guidance in the final investigation. We evaluate FirmRCA with both synthetic and real-world targets, including 41 crashing test cases across 17 firmware images. The results show that FirmRCA can effectively (92.7% success rate) identify the root cause of crashing test cases within the top 10 instructions.

cs.CR

Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference

The prevalent use of Transformer-like models, exemplified by ChatGPT in modern language processing applications, underscores the critical need for enabling private inference essential for many cloud-based services reliant on such models. However, current privacy-preserving frameworks impose significant communication burden, especially for non-linear computation in Transformer model. In this paper, we introduce a novel plug-in method Comet to effectively reduce the communication cost without compromising the inference performance. We second introduce an efficient approximation method to eliminate the heavy communication in finding good initial approximation. We evaluate our Comet on Bert and RoBERTa models with GLUE benchmark datasets, showing up to 3.9$\times$ less communication and 3.5$\times$ speedups while keep competitive model performance compared to the prior art.

cs.LG