SearcharxivSearch

arXiv subjects

Tian Wu

Publications and source records attributed to Tian Wu.

At least 19 recordsLinked to original sources

Subcritical limits of a 1D KPZ growth surface with finite moments

We consider the one-dimensional random growth surface introduced by Adhikari and Chatterjee [AC24], where the height function is defined recursively via the heights at two neighboring sites and an independent noise term at each space-time point. In the subcritical regime, assuming uniformly bounded eighth moments of the noise variables and appropriate regularity of the driving function, we establish convergence in law of the suitably rescaled height function to the solution of the additive stochastic heat equation. Our approach employs a direct recursive renormalization scheme. A central component is a graph representation of the renormalization terms, which enables the derivation of the high-moment bounds essential to the renormalization procedure.

math.PR

Classification of Solutions to a Critical Fourth-Order Equation on Complete Non-compact Manifolds

We study the critical biharmonic equation \[ \Delta^2 u=u^{\frac{n+4}{n-4}}, \] on a complete, connected, and non-compact Riemannian manifold \((M^n,g)\) of dimension \(n\geq 5\) with non-negative Ricci curvature. We establish an optimal pointwise second-order derivative estimate by Bernstein's method and the continuity method. Using invariant tensors, we derive a differential inequality that yields a rigidity result. More precisely, if there exists a positive finite energy solution, then the manifold is isometric to the Euclidean space \(\mathbb{R}^n\), and the solution is given by \[ u(x)= \frac{ \left[\lambda^2 n(n-4)(n^2-4)\right]^{\frac{n-4}{8}} } {\left(1+\lambda |x-x_0|^2\right)^{\frac{n-4}{2}}}, \qquad \lambda>0,\quad x_0\in\mathbb{R}^n . \]

math.DG

Spectral Improvements of Geometric Inequalities on Closed K\"ahler Manifolds

Let $(M,g,J)$ be a closed K\"ahler manifold satisfying $\operatorname{Ric}\geqslant g$. We establish improved Liouville theorems for the Euler--Lagrange equations associated with the Beckner--Sobolev inequalities by incorporating the first positive eigenvalue of the $\bar\partial$-Laplacian into a differential-identity argument. As a consequence, we obtain improved Sobolev and Beckner inequalities that refine the known Riemannian and K\"ahler estimates when the first eigenvalue is sufficiently large. We also derive new upper bounds for the diameter of $(M,g)$.

math.DG

Liouville theorem for a class of p-Laplace type equations on manifolds

We study a class of $p$-Laplace equations $$\Delta_p u-\lambda u^{p-1}+ u^{q-1}=0$$ on a closed $n$-dimensional Riemannian manifold $(M,g)$ with $\operatorname{Ric}\geqslant(n-1)g$. For $1 2$ and $p 0$; aside from the constant solution, the equation admits a positive nonconstant solution. This answers V\'eron's problem raised in \cite{Ver92}.

math.AP

Liouville Rigidity and Universal Spacelikeness Estimates for a Lorentzian Prescribed Mean Curvature Equation

We prove a Liouville theorem for nonnegative entire strictly spacelike solutions of \[ \operatorname{div}\left(\frac{\nabla u}{\sqrt{1-|\nabla u|^2}}\right)+u^p=0 \qquad\text{in }\mathbb R^n. \] If $n=2$ and $p\geqslant1$, or if $n\geqslant3$ and $1\leqslant p<\frac{n+2}{n-2}$, every nonnegative $C^2$ solution satisfying $|\nabla u|<1$ vanishes identically. No symmetry, decay, integrability, or uniform spacelike gap is assumed. A key independent ingredient is a universal bound, valid for every $n\geqslant2$ and $p\geqslant1$, for both the height $u$ and the Lorentz factor $(1-|\nabla u|^2)^{-1/2}$. Thus pointwise strict spacelikeness automatically improves to a uniform spacelike gap, including in the critical and supercritical regimes. The proof combines a geometric Bernstein estimate, comparison with an explicit hyperbolic cap, and weighted trace-free tensor identities. The result extends the known radial nonexistence theorem to arbitrary entire solutions and yields a geometric half-space rigidity theorem for complete spacelike hypersurfaces.

math.AP

Domain-Adaptive Model Merging Across Disconnected Modes

Learning across domains is challenging when data cannot be centralized due to privacy or heterogeneity, which limits the ability to train a single comprehensive model. Model merging provides an appealing alternative by consolidating knowledge from multiple specialized models into one, avoiding data sharing and reducing retraining cost. In this work, we present DMM, a data-free model merging framework designed to handle highly divergent models. DMM proceeds in three steps. First, domain-specific models are trained independently. Second, models with high similarity are merged using standard techniques to ensure stability. Third, we synthesize pseudo-data from normalization statistics and distill knowledge from divergent models into the merged model through a lightweight refinement guided by these samples. This approach preserves rare but critical knowledge while maintaining stability. Extensive experiments on unimodal and multimodal benchmarks show that DMM achieves state-of-the-art performance over existing merging methods.

cs.DC

ERNIE 5.0 Technical Report

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts (MoE) architecture with modality-agnostic expert routing. To address practical challenges in large-scale deployment under diverse resource constraints, ERNIE 5.0 adopts a novel elastic training paradigm. Within a single pre-training run, the model learns a family of sub-models with varying depths, expert capacities, and routing sparsity, enabling flexible trade-offs among performance, model size, and inference latency in memory- or time-constrained scenarios. Moreover, we systematically address the challenges of scaling reinforcement learning to unified foundation models, thereby guaranteeing efficient and stable post-training under ultra-sparse MoE architectures and diverse multimodal settings. Extensive experiments demonstrate that ERNIE 5.0 achieves strong and balanced performance across multiple modalities. To the best of our knowledge, among publicly disclosed models, ERNIE 5.0 represents the first production-scale realization of a trillion-parameter unified autoregressive model that supports both multimodal understanding and generation. To facilitate further research, we present detailed visualizations of modality-agnostic expert routing in the unified model, alongside comprehensive empirical analysis of elastic training, aiming to offer profound insights to the community.

cs.CL

Classification of positive solutions to a class of Laplace equations with a gradient term

In this paper, we investigate positive solutions to a class of Laplace equations with a gradient term on a complete, connected, and noncompact Riemannian manifold \((M^n,g)\) with nonnegative Ricci curvature, namely \[-\Delta u = f(u)|\nabla u|^q\quad\text{in }~M^n,\] where \(n\geqslant 3\), \(q>0,\) and \(f\) is a positive continuous function. We prove some Liouville theorems employing a key differential identity derived via the invariant tensor technique. In particular, for \(f(u)=u^{\frac{2-q}{n-2}(n+\frac{q}{1-q})-1}\) is the second critical case in dimension \(n=3,4,5\), without any additional conditions, such as integrable conditions on \(u\), we show the rigidity for the ambient manifold and classification result of positive solutions. To our knowledge, this is the first rigidity result for equations with gradient terms in the second critical case. Moreover, this result confirms that all solutions must be of the form found in \cite{BV-GH-V2019}.

math.AP

SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference

The Mixture-of-Experts (MoE) architecture has been widely adopted in large language models (LLMs) to reduce computation cost through model sparsity. Employing speculative decoding (SD) can further accelerate MoE inference by drafting multiple tokens per step and verifying them in parallel. However, combining MoE with SD inflates GPU memory and aggravates CPU-GPU bandwidth contention during multi-token verification. Existing MoE offloading systems are SD-agnostic and do not address this bottleneck. We present SP-MoE, the first SD-aware expert-offloading and compute-communication pipelining framework. SP-MoE introduces: (1) speculative expert prefetching that exploits structural correspondence between the draft and target models to prefetch likely experts ahead of verification; (2) a cutoff-layer policy that bounds per-layer prefetch depth based on empirical profiles and an analytical latency model, guaranteeing just-in-time availability without overfetch; and (3) a pipelined runtime with asynchronous prefetch threads and batched I/O to hide loading latency. Extensive experiments demonstrate that SP-MoE achieves a 1.07-3.5 times TPOT speedup over state-of-the-art methods across diverse datasets, environments, and MoE-based models.

cs.DC

The Liouville-type equation and an Onofri-type inequality on closed 4-manifolds

In this paper, we study the Liouville-type equation \[\Delta ^2 u-\lambda_1\kappa\Delta u+\lambda_2\kappa^2(1-\mathrm e^{4u})=0\] on a closed Riemannian manifold \((M^4,g)\) with \(\operatorname{Ric}\geqslant 3\kappa g\) and \(\kappa>0\). Using the method of invariant tensors, we derive a differential identity to classify solutions within certain ranges of the parameters \(\lambda_1,\lambda_2\). A key step in our proof is a second-order derivative estimate, which is established via the continuity method. As an application of the classification results, we derive an Onofri-type inequality on the 4-sphere and prove its rigidity.

math.AP

Liouville theorem of the subcritical biharmonic equation on complete manifolds

In this paper, we study the subcritical biharmonic equation \[\Delta ^2 u=u^\alpha\] on a complete, connected, and non-compact Riemannian manifold $(M^n,g)$ with nonnegative Ricci curvature. Using the method of invariant tensors, we derive a differential identity to obtain a Liouville theorem, i.e., there is no positive $C^4$ solution if $n\geqslant5$ and $1<\alpha<\frac{n+4}{n-4}$. We establish a crucial second-order derivative estimate, which is established via Bernstein's technique and the continuity method.

math.AP

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

The emergence of Mixture-of-Experts (MoE) has transformed the scaling of large language models by enabling vast model capacity through sparse activation. Yet, converting these performance gains into practical edge deployment remains difficult, as the massive memory footprint and communication demands often overwhelm resource-limited environments. While centralized cloud-based solutions are available, they are frequently plagued by prohibitive infrastructure costs, latency issues, and privacy concerns. Moreover, existing edge-oriented optimizations largely overlook the complexities of heterogeneous hardware, focusing instead on isolated or uniform device setups. In response, this paper proposes Prism, an inference framework engineered for collaborative MoE serving across diverse GPU-equipped edge servers. By leveraging the intrinsic sparsity and input locality of MoE workloads, Prism minimizes inter-server communication and optimizes expert placement within diverse resource constraints. The framework integrates an activation-aware placement strategy that balances local request coverage with memory utilization, supplemented by a runtime migration mechanism to adapt expert distribution to dynamic workload changes. Experiments on contemporary MoE models and datasets demonstrate that Prism reduces inference latency by up to 30.6% and significantly lowers communication costs compared to state-of-the-art baselines, confirming the effectiveness of cooperative edge-based MoE serving.

cs.DC

Liouville theorems for anisotropic $p$-Laplace equations with a semilinear term

In this paper, we investigate Liouville theorems for solutions to the anisotropic $p$-Laplace equation $$-\Delta_p^H u=-\operatorname{div}(a(\nabla u))=f(u),\quad\text{in }\mathbb{R}^n,$$ where the semilinear term $f$ may be positive, negative, or sign-changing. When $f$ is positive (negative) and satisfies certain conditions, Serrin's technique is applied to show that every positive supersolution (subsolution) must be constant. For the subcritical case, we use the invariant tensor method to prove nonexistence results for positive solutions. In particular, by applying the differential identity established in the subcritical case to the critical case, we provide a simplified new proof of the classification of positive solutions to the critical case $f(u)=u^{p^*-1}$. For sign-changing solutions, every stable solution or solution that is stable outside a compact set is trivial under certain conditions on $f$.

math.AP

PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning

Fine-grained image classification has witnessed significant advancements with the advent of deep learning and computer vision technologies. However, the scarcity of detailed annotations remains a major challenge, especially in scenarios where obtaining high-quality labeled data is costly or time-consuming. To address this limitation, we introduce Precision-Enhanced Pseudo-Labeling(PEPL) approach specifically designed for fine-grained image classification within a semi-supervised learning framework. Our method leverages the abundance of unlabeled data by generating high-quality pseudo-labels that are progressively refined through two key phases: initial pseudo-label generation and semantic-mixed pseudo-label generation. These phases utilize Class Activation Maps (CAMs) to accurately estimate the semantic content and generate refined labels that capture the essential details necessary for fine-grained classification. By focusing on semantic-level information, our approach effectively addresses the limitations of standard data augmentation and image-mixing techniques in preserving critical fine-grained features. We achieve state-of-the-art performance on benchmark datasets, demonstrating significant improvements over existing semi-supervised strategies, with notable boosts in accuracy and robustness.

cs.CV

Banishing LLM Hallucinations Requires Rethinking Generalization

Despite their powerful chat, coding, and reasoning abilities, Large Language Models (LLMs) frequently hallucinate. Conventional wisdom suggests that hallucinations are a consequence of a balance between creativity and factuality, which can be mitigated, but not eliminated, by grounding the LLM in external knowledge sources. Through extensive systematic experiments, we show that these traditional approaches fail to explain why LLMs hallucinate in practice. Specifically, we show that LLMs augmented with a massive Mixture of Memory Experts (MoME) can easily memorize large datasets of random numbers. We corroborate these experimental findings with a theoretical construction showing that simple neural networks trained to predict the next token hallucinate when the training loss is above a threshold as it usually does in practice when training on internet scale data. We interpret our findings by comparing against traditional retrieval methods for mitigating hallucinations. We use our findings to design a first generation model for removing hallucinations -- Lamini-1 -- that stores facts in a massive mixture of millions of memory experts that are retrieved dynamically.

cs.CL

Classification of positive solutions of critical anisotropic Sobolev equation without the finite volume constraint

In this paper, we classify all positive solutions of the critical anisotropic Sobolev equation \begin{equation}\label{0.1} -\Delta^{H}_{p}u = u^{p^{*}-1}, \ \ x\in \mathbb{R}^n \end{equation} without the finite volume constraint for $n \geq 3$ and $p_n(\Lambda) < p < n$, where $p^{*} = \frac{np}{n-p}$ denotes the critical Sobolev exponent, $-\Delta^{H}_{p}=-div(H^{p-1}(\cdot)\nabla H(\cdot))$ denotes the anisotropic $p$-Laplace operator and $\Lambda = \lambda\max\limits_{\substack{\xi \in \mathbb{R}^n\\1 \leq i, j \leq n}}\left\{\frac{|\xi|^{2}(\nabla^{2}_{ij}H^{p}(\xi))} {p(p-1)H^{p}(\xi)}\right\}$. By employing a novel approach based on invariant tensors technique, and using a Kato-type inequality, we prove that the positive solutions of \eqref{0.1} can be classified for $p_n(\Lambda) \leq p < n$, where $p_n(\Lambda)$ depends explicitly on $\Lambda$. This result removes the finite volume assumption on the classification of critical anisotropic $p$-Laplace equation which was obtained by Ciraolo-Figalli-Roncoroni in the literature \cite{CFR}. In particular, this results capture the precise dependence of critical exponents $p$ on both $n$ and $\Lambda$.

math.AP

Jerison-Lee identity and Semi-linear subelliptic equation on CR manifold

In the study of the extremal for Sobolev inequality on the Heisenberg group and the Cauchy-Riemann(CR) Yamabe problem, Jerison-Lee found a three-dimensional family of differential identities for critical exponent subelliptic equation on Heisenberg group $\mathbb H^n$ by using the computer in [11]. They wanted to know whether there is a theoretical framework that would predict the existence and the structure of such formulae. With the help of dimensional conservation and invariant tensors, we can answer the above question. For a class of subcritical exponent subelliptic equations on the CR manifold, several new types of differential identities are found. Then we use those identities to get the rigidity result, where rigidity means that subelliptic equations have no other solution than some constant at least when parameters are in a certain range. The rigidity result also deduces the sharp Folland-Stein inequality on closed CR manifolds.

math.AP

Documentation based Semantic-Aware Log Parsing

With the recent advances of deep learning techniques, there are rapidly growing interests in applying machine learning to log data. As a fundamental part of log analytics, accurate log parsing that transforms raw logs to structured events is critical for subsequent machine learning and data mining tasks. Previous approaches either analyze the source code for parsing or are data-driven such as text clustering. They largely neglect to exploit another widely available and valuable resource, software documentation that provides detailed explanations for the messages, to improve accuracy. In this paper, we propose an approach and system framework to use documentation knowledge for log parsing. With parameter value identification, it not only can improve the parsing accuracy for documented messages but also for undocumented messages. In addition, it can discover the linkages between event templates that are established by sharing parameters and indicate the correlation of the event context.

cs.SE