SearcharxivSearch

arXiv subjects

Zhuang Tao

Publications and source records attributed to Zhuang Tao.

2 recordsLinked to original sources

OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency

Large Language Models drive a wide range of modern AI applications but impose substantial challenges on large-scale serving systems due to intensive computation, strict latency constraints, and throughput bottlenecks. We introduce OmniInfer, a unified system-level acceleration framework designed to maximize end-to-end serving efficiency through fine-grained optimization of expert placement, cache compression, and scheduling. OmniInfer integrates three complementary components: OmniPlacement for load-aware Mixture-of-Experts scheduling, OmniAttn for sparse attention acceleration, and OmniProxy for disaggregation-aware request scheduling. Built atop vLLM, OmniInfer delivers system-wide performance gains through adaptive resource disaggregation, efficient sparsity exploitation, and global coordination across prefill and decode phases. Evaluated on DeepSeek-R1 within a 10-node Ascend 910C cluster, OmniInfer achieves 616 QPM, where the unified framework reduces TPOT by 36\%, and the superimposition of OmniProxy further slashes TTFT by 38\%. The project is open-sourced at [this https URL](https://gitee.com/omniai/omniinfer).

cs.DC

Newhouse Laminations of polynomials on $\mathbb{C}^2$

It has been recently discovered that in smooth unfoldings of maps with a rank-one homoclinic tangency there are codimension two laminations of maps with infinitely many sinks. Indeed, these laminations, called Newhouse laminations, occur also in the holomorphic context. In the space of polynomials of $\mathbb{C}^2$, with bounded degree, there are Newhouse laminations.

math.DS