SearcharxivSearch

arXiv subjects

Ao Xiao

Publications and source records attributed to Ao Xiao.

3 recordsLinked to original sources

Positive Solutions for a One-Dimensional Minkowski-Curvature Equation with a Sign-Changing Nonlinearity under Mixed Boundary Conditions

We establish the existence of a positive solution for a one-dimensional Minkowski-curvature equation with mixed boundary conditions and a sign-changing nonlinearity. A nonnegative linear shift and the Green kernel of the associated linear problem produce an invariant cone. Principal-eigenvalue comparisons and the fixed point index give cone expansion near zero and compression at infinity. A globally defined auxiliary equation and a first-contact argument show that the resulting solution has slope below one and hence solves the original equation. We also derive a directly verifiable asymptotic criterion.

math.AP

An Infinitesimal Circular Morera Theorem

We prove an infinitesimal circular version of Morera's theorem. Let $D\subset\mathbb{C}$ be a domain and let $f\in C(D)$. If, at every $a\in D$, $\int_{\vert{}\zeta-a\vert{}=r}f(\zeta)\,d\zeta=o(r^2)$ as $r\to0^+$, then $f$ is holomorphic in $D$. In particular, exact vanishing of all sufficiently small centered circular integrals implies holomorphicity. The proof uses a local distributional $\partial$-primitive, a circular identity for weak $\partial$-derivatives, and a pointwise asymptotic mean-value criterion for harmonicity.

math.CV

Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod

Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentralized serving. This report presents xDeepServe, the production serving system behind Huawei Cloud's MaaS offering on CloudMatrix384, a 48-server SuperPod with 384 Ascend 910C chips connected by a high-bandwidth UB fabric and global shared memory. It serves models including DeepSeek, Kimi, GLM, Qwen, and MiniMax, among others. xDeepServe is built around Transformerless, a disaggregated execution architecture that decomposes transformer inference into modular units -- attention, feedforward, and MoE -- and supports disaggregated Prefill-Decode and MoE-Attention deployments. To enable disaggregation, we develop XCCL, a memory-semantic communication layer providing microsecond-level point-to-point and scalable all-to-all primitives, and we extend FlowServe with decentralized DP groups and techniques to mitigate stragglers and synchronization variance. In a peak decoding configuration, xDeepServe reaches 2400 tokens/s per Ascend 910C chip at ~50ms time-per-output-token (TPOT).

cs.DC