SearcharxivSearch

arXiv subjects

Jie Xiao

Publications and source records attributed to Jie Xiao.

At least 19 recordsLinked to original sources

Capacitary-Distance Hardy Inequality

Let $n\ge3$, $\Omega\subset\mathbb R^n$ be an open set, $F:=\mathbb R^n\setminus\Omega$, and $\alpha\in(0,\infty)$. For any $x\in\Omega$, we define the capacitary distance \begin{align*} d_\alpha(x) := \inf\left\{ r>0: \operatorname{cap}(\overline{F\cap B(x,r)}) \ge \alpha\operatorname{cap}(B(\mathbf0,r)) \right\}. \end{align*} In this article, we prove that there exists a positive constant $C_n$, depending only on $n$, such that, for any $\alpha\in(0,1]$ and any $u\in C_{\rm{c}}^\infty(\Omega)$, \begin{align*} \int_\Omega \frac{|u(x)|^2}{d_\alpha(x)^2}\,d x \le \frac{C_n}{\alpha^{2}} \int_\Omega|\nabla u(x)|^2\,d x. \end{align*} This gives an affirmative answer to Problem 8 of Maz'ya [25]. Moreover, this dependence on $\alpha$ is sharp: there exists a positive constant $c_n$, depending only on $n$, such that, for every $\alpha\in(0,1]$, we are able to construct a bounded connected domain $\Omega_\alpha$ on which the optimal constant in the above Hardy inequality is at least $\frac{c_n}{\alpha^{2}}$. The proof combines a variable-time semigroup estimate for the killed Brownian motion with finite-time exit estimates derived from capacity.

math.CA

Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting

In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existing deep forecasting models usually compress them into a unified end-to-end mapping, leading to suboptimal modeling of time-varying dynamics and limited interpretability regarding which forecasting mechanism is activated under different latent states. To overcome these limitations, we reformulate time series forecasting as a unified framework of latent temporal state identification and interpretable expert routing, and propose Fuzzy-MoE, a fuzzy logic-based dynamic Mixture-of-Experts model. Fuzzy-MoE consists of multiple parallel expert mapping networks and a dual-view fuzzy router. By jointly exploiting local convolutional dynamics and global segmented statistics, the router infers latent temporal states and computes expert activation strengths through learnable Gaussian membership functions, enabling explicit IF-THEN rule-based expert selection. This fine-grained routing strategy allows different variables within the same sequence to activate different experts, effectively capturing heterogeneous temporal dynamics while improving model interpretability. Experimental results on multiple public time series benchmark datasets show that Fuzzy-MoE significantly outperforms mainstream forecasting methods in forecasting accuracy. Moreover, fuzzy memberships and rule activations provide interpretable routing diagnostics, demonstrating the effectiveness of the proposed framework in both forecasting performance and mechanism transparency. Unlike traditional MoE models that use black-box routing, Fuzzy-MoE`s routing is based on clear, interpretable fuzzy rules. This makes the expert selection transparent and traceable.

cs.LG

Sharp $p$-Capacity Estimates via Quermassintegrals in Hyperbolic Space

This paper establishes sharp upper bounds for $p$-capacities $\mathrm{Cap}_{1 2m+1$, an interpolating radius combines the $2m$-th curvature-excess radius with the $L^\infty$ curvature scale, thereby linking the finite-moment and supremum regimes. Equality in the sharp comparisons characterizes geodesic balls.

math.DG

A Unified Quermassintegral Approach to Quasilinear Heat Dispersion and Loss

This paper establishes a fundamental connection between quasilinear potential theory and convex geometric analysis by investigating the interplay between the quasilinear Laplace operator and quermassintegrals. We introduce a quasilinear heat dispersion law for convex conductors and prove that, among all convex conductors of a fixed mean width, the closed ball is a unique maximizer of this dispersion. By characterizing the quasilinear heat loss of a convex conductor explicitly in terms of its quermassintegrals, we demonstrate not only a formal equivalence between the isocapacitary and isoperimetric inequalities in the setting of mathematical physics but also that, among all convex conductors of a fixed mean width, the closed ball is a unique maximizer of this loss. These results provide a novel bridge between the metric properties of convex conductors and the variational analysis of quasilinear elliptic operators, offering a unified perspective on sharp geometric inequalities and their extremal cases.

math.AP

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. Such policies reduce inference cost, but they leave the small model's capability unchanged, so attainable savings remain bounded by the work the student can already solve. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. In each cycle, MERA replays failed student invocations to obtain execution-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine-tunes a student LoRA adapter via supervised learning and optional GRPO. Routing serves as supporting machinery for deployment: the improved student is served behind a cost-calibrated router with verifier-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality. Empirically, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP. Under verifier-backed fallback, the deployed policy retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, a fine-tuned Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B model. These results indicate that verifier-backed multi-cycle adaptation can increase small-model capability, rather than only routing around a fixed student.

cs.LG

A Counterexample to the Liu--Lou--Zhu $\mathcal Q_p$--Carleson Embedding Conjecture

In this paper, we disprove a conjecture of Liu, Lou, and Zhu concerning Carleson embeddings of $\mathcal Q_p$ spaces into tent spaces for $0<p<1$. More precisely, we construct a finite positive $p$-Carleson measure $\mu$ on the unit disc $\mathbb D$ such that the canonical embedding $$ \operatorname{id}:\mathcal Q_p \longrightarrow \mathcal T_{p,2}^2(\mu) $$ is not bounded. The main ingredient is a new family of $\mathcal Q_p$ test functions that encodes Cantor-type structures on the unit circle $\mathbb T$ into the analytic behavior of functions in $\mathcal Q_p$. This construction is inspired by ideas developed in a recent work of the first and third authors on composition operators on $\mathcal Q_p$ spaces.

math.CV

On the holomorphic differential operator $\frac{d}{dz}: Q_K\to L^q(WdA)$

In this paper, we obtain non-testing characterizations, in terms of dyadic capacity gauges, of the boundedness and compactness of the differentiation operator $$ \frac{d}{dz}:Q_K\longrightarrow L^q(W\,dA), \qquad 0<q<\infty. $$ We also characterize the limiting case as $q\to0^+$, formulated in terms of a logarithmic geometric mean, while the endpoint $q=\infty$ is treated separately using a standard testing argument. These results greatly extend the previous work on ${\mathcal Q}_p$-spaces to the general setting of $Q_K$-spaces. As applications, we characterize composition operators and Volterra-type integral operators between different $Q_K$-spaces. In particular, the off-diagonal characterization established here, together with the previously established diagonal case, completely resolves Zhao's 2009 open question on composition operators between ${\mathcal Q}_p$-spaces.

math.CV

On the growth of Bloch functions

We prove that there exist two Bloch functions $f_1$ and $f_2$ on $\mathbb D$ such that $$ |f_1(z)|+|f_2(z)| \geq \left(\log\frac{1}{1-|z|}\right)^{1/2}, \qquad z\in\mathbb D, $$ thereby resolving an open problem posed in 2008 by Girela, Pel\'aez, P\'erez-Gonz\'alez and R\"atty\"a. Our proof is based on a new Szeg\H{o}-type recursion involving $\mathbb C^2$-valued polynomials and their reciprocal polynomials.

math.CV

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population objective. Under assumptions of local boundedness, distributional smoothness, and behavior-policy smoothness, we show that stale rollouts introduce a per-step surrogate-gradient bias of order O(S * eta), where S denotes the maximum rollout lag and eta denotes the learning rate. We further derive a conditional collapse-time scaling law: when within-cycle drift remains below a batch-level clipping radius, collapse is governed primarily by cumulative learner drift T * eta; when the stale-rollout constraint is active, stability instead depends explicitly on S * eta. This yields a two-constraint stability condition eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)}, explaining why the maximum stable learning rate may appear weakly dependent on staleness in the horizon-limited regime.

cs.LG

Proximity-Induced Skyrmion Stabilization at the Cu2OSeO3/Bi2Se3 Interface

We investigate proximity-induced magnetic interactions at the interface between the topological insulator Bi2Se3 and the chiral magnetic insulator Cu2OSeO3, with particular focus on the low temperature skyrmion phase. Broadband ferromagnetic resonance spectroscopy reveals enhanced stability of noncollinear spin textures in the Cu2OSeO3/Bi2Se3 heterostructure compared with bare Cu2OSeO3. In addition to an extra resonance mode in the tilted conical phase that is absent in bare Cu2OSeO3, field cycling resolves two counterclockwise skyrmion resonance branches separated by approximately 238 MHz, consistent with the coexistence of a bulk skyrmion lattice and an interfacial skyrmion phase stabilized by proximity-induced exchange coupling and enhanced interfacial Dzyaloshinskii-Moriya interactions. The finite frequency separation indicates that the two skyrmion phases occupy distinct magnetic energy landscapes while retaining similar resonance character. Resonant elastic x-ray scattering measurements further confirm that the interfacial skyrmion phase spans a broader magnetic-field range than the bulk phase, demonstrating enhanced stability and ordering of topological spin textures at the interface. These findings establish interface engineering as a promising route for extending the stability regime of skyrmion and tilted-conical phases in topological-magnetic heterostructures.

cond-mat.mtrl-sci

Geometric realization of affine bases: the Kronecker quiver case

In this paper, we study the transition matrix between the PBW basis and the canonical basis for the negative part of the quantized enveloping algebra of the Kronecker quiver from a geometric viewpoint. Building on Lusztig's geometric construction of the canonical basis, we construct sheaf-complex realizations of PBW basis elements by means of flag sheaf complexes over the strata $X(\alpha,m)$ of representation varieties. Our first goal is to give a geometric description of the simple constituents appearing in the restrictions of these flag sheaf complexes to the strata $X(\alpha,m)$. This allows us to compare the PBW-type sheaf complexes with the simple perverse sheaves $IC(X(\alpha),L_\chi)$ arising in Lusztig's construction. Using this description together with a purity result for the relevant $\mathbb{F}_q$-structures, we obtain another proof that the elements defined by Lusztig's perverse sheaves indeed form a basis of the composition algebra.Our second goal is to make the transition coefficients between the PBW basis and the canonical basis geometrically explicit. More precisely, we show that these coefficients are governed by the multiplicities of local systems in the restrictions of intersection cohomology complexes to smaller strata. As a consequence, the transition matrix from the canonical basis to the PBW basis is upper triangular with diagonal entries equal to $1$, and its coefficients admit a direct geometric interpretation. In particular, in the Kronecker quiver case we recover the triangularity of the transition matrix and obtain positivity properties of the corresponding coefficient polynomials.

math.QA

The iterated geometric Green's formula

Fang, Lan, and Xiao established the geometric Green's formula as a categorical isomorphism for arbitrary semisimple complexes. In this short note, we generalize their work to multi-step compositions. Specifically, we establish the iterated geometric Green's formulas for the composition of an $(n-1)$-fold restriction and an induction, as well as its dual.

math.RT

T2S: A Rehearsal-Based Approach for Extraction-Resistant Model Watermarking

Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks emerge as the most severe threat, where adversaries exploit prediction outputs to train surrogate models that illegally replicate the original model's functionality. In this work, we propose a rehearsal-based watermark embedding framework to enhance the robustness of model watermarks against model extraction attacks. By simulating the extraction process, our method leverages the loss of a \textit{simulated stolen model} on a trigger set as a training signal to fine-tune the watermark knowledge within the target model. This fine-tuning step encourages the watermark to be embedded in a way that boosts transferability, thereby increasing its chances of persisting and remaining detectable in stolen models. Comprehensive experiments conducted under diverse settings demonstrate that the proposed method significantly improves the robustness of model watermarks against both model extraction and subsequent watermark removal attacks.

cs.CR

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constrained by a dual misalignment between authentic generation trajectory and the gradient update process: (i) Process-reward misalignment. Sparse, terminal rewards are indiscriminately assigned to all intermediate steps of the generation process, failing to provide discriminative credit assignment. (ii) State-trajectory misalignment. Policy updates are often diverted toward artificial, out-of-trajectory states, squandering gradients on less informative samples. To address these limitations, we introduce Process Aligned Policy Optimization (PAPO), a novel framework that holistically aligns the RL update with the dLLM's generative trajectory via Step-Aware Process Rewards (SPR) that transform sparse terminal rewards into dense, step-wise credit, and Entropy-Guided Historical Re-enactment (EHR) that replays authentic trajectories at high-uncertainty steps. Extensive experiments on four benchmarks demonstrate that PAPO significantly outperforms baselines, achieving gains of 4.5% on GSM8K, 4.8% on MATH500, 42.2% on Countdown and 16.1% on Sudoku.

cs.CL

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge

Automated mitosis detection is a well-established task in computational pathology. While previous benchmarks focused on scanner-induced domain shift, clinical "real-world" application requires models to be robust across the vast variance to be expected in the histological landscape. The MItosis DOmain Generalization (MIDOG) 2025 challenge was designed to evaluate algorithmic performance across unprecedented biological and contextual diversity. We curated a test dataset of 365 cases, encompassing 12 distinct human, canine and feline tumor types, digitized across multiple scanning platforms. Moving beyond hand-selected hotspots, the challenge required detection also in random tissue areas (representative of the whole slide detection situation) and challenging areas (areas rich in hard negatives). In the second track, we introduced the classification of atypical mitotic figures (AMFs). There were 18 teams submitting to the detection track, with F1 scores ranging up to 0.740. In the AMF detection track, we had 21 submissions with balanced accuracy values up to 0.908. Our analysis reveals that while most models perform reliably in traditional hotspots, significant performance degradation occurs in challenging ROIs, where false positive rates tripled. Furthermore, performance varied significantly across the 12 tumor types, highlighting "blind spots" in current state-of-the-art architectures when encountering rare or highly pleomorphic malignancies. Moreover, we evaluated the effectiveness of ensembling and found a mean increases of 1.5 and 1.3 percentage points in F1 score and balanced accuracy, respectively. In contrast, TTA showed no relevant improvement. MIDOG 2025 demonstrates that "in the wild" mitosis detection remains a significant hurdle. The transition from hotspot-only evaluation to a multi-contextual framework provides a more realistic proxy for clinical reliability.

cs.CV

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety control fundamentally different from prompt-level filtering or output-level detection. Harmful semantics may be weakly expressed in text representations, progressively bound to visual latents, and finally entangled with rendering dynamics. As a result, safety steering at a fixed layer can be unstable, and a steering mechanism learned from known risks may not transfer reliably to a shifted target risk domain. We propose SafeDIG, a safety steering framework that formulates DiT safety adaptation as position-aware sparse feature transfer. SafeDIG first constructs Sparse Autoencoders over functionally distinct DiT intervention positions and uses robustness-aware pre-training routing to prioritize intervention sites that are expected to remain stable under source-target risk shift. It then separates transferable safety features from domain-specific activation geometry by freezing the SAE encoder as a reusable sparse safety dictionary and adapting only the decoder to the target-domain activation manifold. During inference, SafeDIG combines Blend and Repel operations to steer unsafe activations toward transferred safety manifolds or away from harmful sparse directions. Experiments on FLUX.1 Dev and Stable Diffusion 3.5 Large show that SafeDIG consistently reduces target-domain and overall unsafe generation rates while preserving source-domain safety and image quality.

cs.AI