SearcharxivSearch

arXiv subjects

Lei Zhang

Publications and source records attributed to Lei Zhang.

At least 19 recordsLinked to original sources

Rankin--Selberg integrals of opposite conductor--one newforms

Let $F$ be a nonarchimedean local field of characteristic zero and let $n\geq2$. For $r=n,n+1$, let $\Pi_r$ be an irreducible tempered representation of ${\rm GL}_r(F)$ of conductor one and with trivial central character. We evaluate the Rankin--Selberg integral of opposite newforms in $\Pi_{n+1}\times \Pi_n$ explicitly and show that its central value is nonzero. As an application, this implies a case of Disegni--Zhang's conjecture on the nonvanishing of local relative characters.

math.RT

Logarithmic Stability for an Inverse Source Problem in a Coupled Nonlinear Helmholtz System

We study an inverse source problem for a two-mode nonlinear Helmholtz system motivated by second-harmonic generation. The datum is the full first Fr\'{e}chet derivative of the nonlinear Dirichlet-to-Neumann map at one fixed, common, small boundary state, which need not be zero. For real-valued small data, a known susceptibility bounded away from zero, and sources supported in a fixed compact subset of the domain, we prove uniqueness and a conditional single-logarithmic stability estimate. The first boundary variation is the Dirichlet-to-Neumann map of a symmetric $2\times2$ matrix Schr\"odinger operator. Absorbing the two known Helmholtz energies into its matrix potential permits the use of standard zero-energy complex geometrical optics solutions with a common null phase geometry. A bilinear Alessandrini identity then gives a uniform Fourier estimate for the two entries containing the background fields. Zero extension below the exponent $3/2$, a low/high frequency splitting, and interpolation between $H^{-2}$ and the a priori $H^s$ source bound yield the stated stability modulus.

math.AP

TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents

GUI agents accumulate high-resolution screenshots as the trajectory unfolds, increasing inference latency and memory usage. Training-free visual token pruning can reduce this cost, but cache reuse introduces a fundamental constraint. Once tokens are discarded, the corresponding visual evidence cannot be recovered without re-encoding. Pruning therefore becomes an \textit{irreversible admission decision} that must remain useful for unknown future targets while preserving coverage of operable regions under tight budgets. To address these challenges, we propose \textbf{\method{}}, a training-free framework for \emph{\textbf{T}rajectory-\textbf{r}obust \textbf{A}dmission and \textbf{C}overage-aware \textbf{E}vidence ordering}. Specifically, we combine a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity. Then, we reserve part of the budget for native visual tokens distributed across the screen, repairing missing spatial coverage without breaking the ordering. Together, these mechanisms produce a nested token order, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory. Finally, our monotone KV contraction incrementally contracts retired frames into compact session state, avoiding repeated visual encoding or pruning. Extensive experiments across six GUI benchmarks and diverse models verify the effectiveness of our proposed \method{} under tight budgets. The source code will be released.

cs.CV

QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement Learning

Quality-of-service (QoS)-aware random access requires adaptive allocation of a finite random access channel (RACH) preamble budget across heterogeneous traffic and access procedures. This paper proposes QP-BD3QN-RACH, a quota-projected branching deep reinforcement learning controller for mixed two-step (2RA) and four-step (4RA) contention-based random access. Four action branches correspond to the delay-sensitive and delay-tolerant 2RA/4RA preamble pools. A branching dueling Double DQN selects pool-specific multipliers, and deterministic quota projection converts them to nonnegative integer allocations that preserve the preamble budget. With five actions per branch, the controller represents 625 pre-projection branch-action tuples using 20 branch-action outputs. Evaluation covers five arrival loads, cross-method comparison under nominal seed 42, six-seed sensitivity of QP-BD3QN-RACH, and targeted ablations. Across the five-load grid, its mean direction-aligned differences relative to four comparators are positive: 5.74 to 8.21 percentage points for success/collision, 1.23 to 1.92 percentage points for fallback, 0.35 to 0.68 percentage points for blocking, and 0.128 to 0.456 decision intervals for successful-access delay. Load-wise results exhibit metric-dependent tradeoffs, particularly under intermediate and overload conditions.

cs.NI

First-Order Optimization under Uniform Nondegeneracy: Geometry, Computation, and Information

Strongly convex minimization and its natural indefinite extension to strongly convex--strongly concave minimax problems combine quantitative control of curvature with a prescribed curvature orientation. We disentangle these two roles by retaining uniform nondegeneracy alone: curvature remains uniformly separated from zero but may have either sign, with no prescribed positive--negative splitting. Surprisingly, a large part of the familiar theory nevertheless re-emerges. We first derive an intrinsic formulation through first-order secant inequalities, making the gradient on $\mathbb{R}^d$ a global bi-Lipschitz homeomorphism and yielding a unique stationary point. We then pair signed Moreau envelopes to construct a smooth scalar merit that recovers the missing descent geometry at both zeroth and first order, and the resulting paired proximal descent method achieves global linear convergence with dimension-free first-order oracle complexity. Meanwhile, we show that this tractability can break down on restricted domains: merely assuming the existence of a stationary point in the domain may lead to the curse of dimensionality, even with access to an infinite-order oracle. To overcome this information barrier, we introduce certified feasibility, an observable localization condition that enables feasible continuation. Together, these results establish a first-order optimization theory under uniform nondegeneracy that spans geometry, computation, and information.

math.OC

Asymptotic properties of the Bingham distribution in arbitrary dimensions

The Bingham distribution is widely used to model directional data with antipodal symmetry, and a central analytical object is the normalizing constant $Z$. Even though efficient algorithms exist for bounded or moderate parameter values, its singular behavior in general dimensions remains underexplored in the literature. We use a well-known integral representation to obtain asymptotic series of $Z$ and its derivatives, which lead to asymptotic profiles of its moments. Then, we also study the asymptotic behavior of the Bingham entropy, and prove its decomposition into an explicit logarithmic leading term and a Lipschitz continuous correction term. The degeneration of higher-dimensional Bingham distributions to their lower-dimensional counterparts is observed in both the asymptotic series and the entropy.

math.PR

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility remains a persistent challenge, particularly in handling diverse regional dialects, dynamic entities and hotwords, long-range contextual information, and disfluent spontaneous speech. In this report, we present Qwen-Audio-3.0-ASR, a Mixture-of-Experts (MoE) LLM-based ASR system designed to address these production demands through a unified, instruction-following framework. The model is built upon the Qwen backbone, and is trained on tens of millions of hours of large-scale speech data. Qwen-Audio-3.0-ASR supports transcription across 30 languages and 16 Chinese dialectal varieties spanning eight major dialect regions. Beyond multilingual and dialectal recognition, the model provides production-oriented capabilities including industry-domain entity recognition, hierarchical hotword customization, native single-pass transcription polishing, and long-audio contextual modeling. We further develop a dedicated streaming variant, Qwen-Audio-3.0-ASR-Streaming, for latency-sensitive applications. Extensive evaluations on Chinese, English, multilingual, and real-world industrial test sets demonstrate state-of-the-art or highly competitive recognition performance across a broad range of evaluation conditions, with strong performance relative to leading commercial and proprietary systems including GPT-4o Transcribe and Gemini 3.1 Pro.

cs.CL

Stability of independence polynomials of spiders

For a graph $G$, let $i_k(G)$ denote the number of independent sets of cardinality $k$, and let \[ I(G,z)=\sum_{k\ge0} i_k(G)z^k \] be its independence polynomial. Following Brown and Cameron \cite{BrownCameron2018}, a graph is called stable if all zeros of its independence polynomial lie in the closed left half-plane. They proved that every star is stable, but also constructed nonstable trees. They then asked for a characterization of stable trees. In this paper, we extend and strengthen their result by proving that every spider, obtained from a star by arbitrary and possibly nonuniform subdivisions of its edges, has all its independence roots in the open left half-plane. Hence, every spider is stable.

math.CO

Simulation-free Unbalanced Dynamic Optimal Transport with General Growth Penalty

Inferring cellular dynamics from unpaired single-cell snapshots requires modeling both state transitions and population growth or death. Unbalanced dynamic optimal transport (UDOT) addresses this by penalizing growth along transport paths, making the choice of growth penalty a key way to encode biological priors on proliferation and apoptosis. However, existing UDOT solvers either rely on computationally expensive NeuralODE simulations or depend on analytical solutions of conditional paths, restricting their efficiency solely to quadratic penalties, i.e. Wasserstein-Fisher-Rao (WFR) geodesics. To enable an efficient UDOT solver for general growth penalties, we first show that concave growth penalties lead to degenerate solutions where growth and transport are separated. We then introduce \textbf{S}imulation-free \textbf{U}nbalanced \textbf{D}ynamic \textbf{O}ptimal transport (SUDO), a simulation-free framework for UDOT with general non-quadratic convex growth penalties. SUDO learns the conditional paths and transport costs, solves the induced semi-coupling problem, and subsequently leverages unbalanced flow matching to achieve a simulation-free solution. On WFR benchmarks, SUDO matches the accuracy of efficient, analytical solution-driven algorithms while outperforming simulation-based methods in computational speed. Beyond WFR, SUDO supports asymmetric penalties that encode proliferation-dominant priors and produce more plausible trajectories and growth estimates on synthetic and single-cell datasets.

cs.LG

Computing stable configurations of confined smectic liquid crystals with a deep variational framework

Smectic liquid crystals are layered liquid-crystalline phases characterized by orientational order and periodic density modulation. Although their structures can be modeled using continuum theories, computing stable configurations remains challenging in complex geometries, particularly when the high-frequency density modulations associated with smectic layering should be resolved. We propose a deep variational framework (DVF) for computing these configurations within the modified Landau--de Gennes model, in which the coupled orientational and positional order parameters are represented on a regular reference domain while physical confinement is incorporated through coordinate mappings. A warmup penalty mitigates the spectral bias of neural networks toward smooth, nonlayered fields, enabling robust recovery of oscillatory smectic states. Comparisons with a neural-network baseline and finite-difference relaxation demonstrate the essential role of this penalty and the numerical stability of the resulting layered states. The DVF reproduces experimentally established smectic-A defect structures and layer morphologies across diverse confinement geometries and further predicts a chevron-like smectic-C state in a tangent-anchored sphere. Together, these results demonstrate the applicability of the DVF to computing stable smectic configurations across experimentally relevant confinement geometries and anchoring conditions.

cond-mat.soft

Stochastic Operator Inference for reduced-order modeling of capillary wave turbulence using experimental measurements

Modeling complex physical phenomena directly from experimental data poses fundamental challenges: measurement devices introduce noise and biases into the data, the governing equations are often unknown or intractable, and the dynamics observed experimentally may exhibit stochastic behavior. In this paper, we consider capillary wave turbulence---an example of nonlinear wave interactions at a microfluidic interface---measured by ultra-high-speed digital holographic microscopy. With the goal to learn an efficient model directly from these data, we adapt a stochastic extension of Operator Inference to learn low-dimensional stochastic differential equation representations of the microscale wave dynamics from experimental measurements. Each experiment is repeated 16 times, and 10 different conditions are created by changing the nondimensional acoustic capillary number through an excitation device. In addition, we propose a new strategy to select the reduced-order model dimension, based on both mean and covariance errors. We further add Tikhonov regularization to the stochastic Operator Inference framework and show that, beyond its conventional role as a numerical stabilization technique in deterministic settings, it has a physically meaningful interpretation as controlling the spectral content of the learned stochastic dynamics. We demonstrate that the resulting stochastic reduced-order models faithfully capture salient physical features of capillary wave turbulence across a range of experimental conditions.

physics.comp-ph

Simulation of the electron escape ratio for keV alpha-particle ionization tracks in liquid helium

The electron escape ratio for 5.3 MeV alpha-particle ionization tracks in liquid helium under varying external electric fields has been measured in several experiments. However, to the best of our knowledge, the corresponding ratio for keV-scale alpha tracks in the same medium has not yet been reported. In this article, we demonstrate for the first time that this ratio can be accurately characterized using COMSOL-based simulations. Our simulation framework was developed in two stages. In Stage I, we aimed to verify consistency between our simulated results and published experimental data for 5.3 MeV alpha particles. Following successful verification in Stage I, we proceeded to Stage II, in which the 5.3 MeV track was replaced by 2, 5, and 10 keV tracks. Our simulated results reveal that (a) keV-scale tracks exhibit electron escape ratios approximately 1.5-2.5 times higher than that of the 5.3 MeV track, and (b) the escape ratios for all track energies (2, 5, 10 keV, and 5.3 MeV) exhibit a linear dependence on the ion number density at the simulation's T0, but not on the electron number density.

astro-ph.IM

Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence

Artificial intelligence is changing the form of applied English materials from fixed paper sequences to adaptive learning systems that can diagnose learners, recommend tasks, and provide formative feedback. This paper studies the structure and application of a new practical English textbook driven by artificial intelligence. A five-layer architecture is proposed: knowledge mapping, learner profiling, task generation, feedback orchestration, and teacher-side governance. A prototype was tested on 186 non-English-major undergraduates for eight weeks of teaching. Compared with a static digital textbook, the proposed system increased the unit completion accuracy from 72.4% to 84.9%, raised the average score for speaking tasks by 10.8 points, and reduced the teacher's correction time by 31.6%. Therefore, an AI-driven textbook can maintain the stability of the curriculum while providing personalised learning paths, rich practice materials and traceable classroom data.

cs.AI

Real-Time Shape Control of Multi-Segment Soft Robotic Arms Using Koopman Operators with Global and Local Observables

Multi-segment soft robotic arms can continuously reconfigure their body shapes for safe interaction, but tip control alone is insufficient for constrained-space tasks. Therefore, shape control is a more important task for multi-segment soft arms than tip control, but remains challenging due to the high dimensionality and nonlinear dynamics of continuum deformation. In existing work, shape control accuracy is defined by the error in the global frame (global shape error). For multi-segment soft arms, using only global shape error as the control objective is insufficient, as segment coupling, gravity-induced loading, and inertial effects become more significant. This difficulty increases with the number of segments. In this paper, we present a Koopman-based model predictive control framework that combines global and local observables, enabling real-time shape control on multi-segment soft robotic arms. The framework is evaluated through numerical and physical experiments. Numerical experiments demonstrate the scalability of the proposed controller by achieving shape control on robots with up to 10 independently actuated segments. The physical experiments demonstrate that the controller is capable of (1) real-time shape control of 3- and 5-segment robotic arms with tip speeds up to 0.6 m/s, (2) robust tracking without retraining, including distal payloads up to 400~g and recovery from a 7~N lateral disturbance, and (3) the potential for future inspection applications through a confined-space demonstration. These results demonstrate that the proposed framework enables dynamic, scalable, and accurate real-time shape control on multi-segment soft robotic arms.

cs.RO

Uniqueness and Stability of Monge--Amp\`ere Potentials in Big Cohomology Classes

In this paper, we establish a stability estimate and a uniform estimate for the modulus of continuity of solutions to the degenerate complex Monge-Amp\`ere equation in big cohomology classes. Consequently, we prove the uniqueness of solutions to complex Monge-Amp\`ere mean field equations for a sufficiently small parameter.

math.DG

Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models

As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.

cs.CV

Seeing the World and the Self from Egocentric Video

Complete 3D perception from egocentric video requires recovering the surrounding scene and the wearer's full-body motion in a shared metric frame. Existing methods typically address scene reconstruction and motion estimation separately: scene reconstruction methods ignore the wearer, whereas motion estimation methods lack explicit scene geometry and often depend on external trajectories. Joint recovery is challenging because the two tasks exhibit asymmetric visibility and require different prediction paradigms. The largely visible scene supports deterministic geometric regression, whereas the severely occluded body requires generative motion inference. We therefore propose RESELF (REconstructing the Scene and the sELF), a unified framework that couples deterministic metric geometry reconstruction with geometry-conditioned motion generation. RESELF adapts a geometry foundation model pre-trained on large-scale exocentric data to egocentric video using frame-wise scale and relative-pose consistency objectives. The resulting camera trajectory and latent geometric features condition a diffusion model that recovers the wearer's motion. A subsequent closed-loop kinematic feedback stage further refines the camera head while preserving the reconstructed scene geometry. To support training and evaluation, we curate EE4D-JSM from EgoExo4D by aligning egocentric video, sparse metric scene geometry, camera trajectories, and full-body motion annotations. Experiments show that RESELF outperforms state-of-the-art methods designed for the individual tasks across depth estimation, camera tracking, and full-body motion estimation. Code, models, and datasets will be available at https://ka1guan.github.io/RESELF/.

cs.CV

ByteX: A Unified AI Search Engine at ByteDance

Since 2016, ByteX has been the foundation of ByteDance's search infrastructure, scaling to more than 7,000 clusters and 300 PB of indexed data. Driven by the demands of AI workloads, ByteX has evolved from a text search engine into a unified AI search system supporting vector retrieval, lexical matching, and predicate filtering. Its largest deployment indexes nearly one trillion high-dimensional vectors. This scale exposes two central bottlenecks in AI-era retrieval: memory-intensive graph-index construction under sustained ingestion, and the prohibitive cost of keeping vector indexes entirely in memory. ByteX addresses these bottlenecks with two techniques. First, it introduces a quantization-aware vector kernel based on SymRaBitQ, a new symmetric quantization scheme with tight theoretical guarantees that allows index construction to run directly in the quantized space accurately and efficiently without retaining a copy of full-precision vectors. Second, it provides a hybrid storage engine that supports memory-resident, hybrid, and SSD-resident deployments, with fine-grained record-level caching to trade memory for latency under operational control. On large-scale benchmarks, ByteX improves throughput by up to 3x, reduces indexing memory by 80%, and lowers operating cost by 86% compared with prior systems, while supporting trillion-vector scale, write-heavy or latency-sensitive workloads in production.

cs.DB