Searcharxiv⌕ Search

arXiv subjects

Ze Tao

Publications and source records attributed to Ze Tao.

20 records · Page 2Linked to original sources

Analytical and Neural Network Approaches for Solving Two-Dimensional Nonlinear Transient Heat Conduction

Accurately predicting nonlinear transient thermal fields in two-dimensional domains is a significant challenge in various engineering fields, where conventional analytical and numerical methods struggle to balance physical fidelity with computational efficiency when dealing with strong material nonlinearities and evolving multiphysics boundary conditions. To address this challenge, we propose a novel cross-disciplinary approach integrating Green's function formulations with adaptive neural operators, enabling a new paradigm for multiphysics thermal analysis. Our methodology combines rigorous analytical derivations with a physics-informed neural architecture consisting of five adaptive hidden layers (64 neurons per layer) that incorporates solutions as physical constraints, optimizing learning rates to balance convergence stability and computational speed. Extensive validation demonstrates superior performance in handling rapid thermal transients and strongly coupled nonlinear responses, which significantly improves computational efficiency while maintaining high agreement with analytical benchmarks across a range of material configurations and boundary conditions.

physics.comp-ph↗

ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition

Self-attention is an essential component of large language models (LLM) but a significant source of inference latency for long sequences. In multi-tenant LLM serving scenarios, the compute and memory operation cost of self-attention can be optimized by using the probability that multiple LLM requests have shared system prompts in prefixes. In this paper, we introduce ChunkAttention, a prefix-aware self-attention module that can detect matching prompt prefixes across multiple requests and share their key/value tensors in memory at runtime to improve the memory utilization of KV cache. This is achieved by breaking monolithic key/value tensors into smaller chunks and structuring them into the auxiliary prefix tree. Consequently, on top of the prefix-tree based KV cache, we design an efficient self-attention kernel, where a two-phase partition algorithm is implemented to improve the data locality during self-attention computation in the presence of shared system prompts. Experiments show that ChunkAttention can speed up the self-attention kernel by 3.2-4.8$\times$ compared to the state-of-the-art implementation, with the length of the system prompt ranging from 1024 to 4096.

cs.LG↗