SearcharxivSearch

arXiv subjects

Meng Cai

Publications and source records attributed to Meng Cai.

At least 19 recordsLinked to original sources

Strong order one-half convergence of a coupled tamed Euler--Peano scheme for reflected stochastic differential equations with super-linearly growing coefficients

We study strong numerical approximations for reflected stochastic differential equations in possibly unbounded convex domains with super-linearly growing drift and diffusion coefficients. Under a coupled monotonicity condition and polynomial local Lipschitz assumptions, we first establish the well-posedness of the reflected SDE and derive uniform moment bounds for its solution. We then introduce a coupled tamed Euler--Peano scheme, in which the drift and the squared diffusion coefficient are tamed by a common factor and the resulting Euler--Peano path is corrected through the Skorokhod problem. This common taming factor preserves the drift--diffusion coercivity structure and yields uniform moment estimates for the numerical solution. We prove strong convergence of order $1/2$ for both the constrained state process and the boundary regulator, thereby recovering the standard Euler-type strong order in this reflected setting. Numerical experiments for a reflected stochastic Ginzburg--Landau type system illustrate the constraint preservation of the scheme and support the theoretical convergence rate.

math.NA

ARBITER: Reasoning Trajectory Basins and Majority Vote Failures in Test-Time Sampling

When language models use test-time sampling, they generate multiple reasoning trajectories and select an answer by majority vote. We show that these trajectories are not independent: for a given question, they concentrate into a small number of clusters, or reasoning basins, each defined by a normalized final answer and the solutions that reach it. A majority vote therefore selects the most stable basin rather than the most accurate one, which creates wrong-majority failures where the correct answer is present but outvoted. We introduce ARBITER, a model-agnostic approach that models interactions between basins using only the base model's own sampled outputs, hidden states, and derived evidence. Most direct correction strategies fail; ARBITER instead uses conservative additive evidence on top of consensus. In its simplest parameter-free form, ARBITER-{\Delta} adds same-model evidence to the majority prior, while ARBITER-Enc augments this with bounded residual signals from hidden states over complete solutions. On GSM8K with Qwen3-4B, consensus over K=24 samples achieves around the mid-94% range, while a same-pool top-2 oracle reaches around the mid-96% range. ARBITER recovers a subset of these cases using zero external information. Across three model families and three math benchmarks, it yields consistent gains with no net-negative cases; for example, on Llama-3.1-8B MMLU-HS-Math, it improves accuracy from the mid-78% range to the mid-82% range, recovering about 22% of the available oracle headroom, indicating that this headroom can be partially recovered from the sample pool itself.

cs.LG

Weak convergence order of stochastic theta method for SDEs driven by time-changed L\'{e}vy noise

This paper studies the weak convergence order of the stochastic theta method for stochastic differential equations (SDEs) driven by time-changed L\'{e}vy noise under global Lipschitz and linear growth conditions. In contrast to classical L\'{e}vy-driven SDEs, the presence of a random time change makes the weak error analysis involve both the discretization error of the underlying equation and the approximation error of the random clock. Moreover, compared with explicit Euler--Maruyama method, the implicit drift correction in the stochastic theta method makes the associated weak error analysis substantially more delicate. To address these difficulties, we first establish a global weak convergence estimate of order one for the stochastic theta method applied to the corresponding non-time-changed L\'{e}vy SDEs on the infinite time interval by means of the Kolmogorov backward partial integro differential equations. Incorporating the approximation of the inverse subordinator together with the duality principle, we derive the weak convergence order of the stochastic theta method with $\theta \in [0,1]$ in the time-changed L\'{e}vy setting. The result advances the currently available weak convergence analysis beyond the Euler--Maruyama method to the more general class of stochastic theta method, and establishes a workable route from the weak analysis of the underlying non-time-changed L\'{e}vy equation to the corresponding time-changed problem. Finally, numerical experiments are presented to further support the theoretical findings.

math.NA

AuxDet: Auxiliary Metadata Matters for Omni-Domain Infrared Small Target Detection

Omni-domain infrared small target detection (Omni-IRSTD) poses formidable challenges, as a single model must seamlessly adapt to diverse imaging systems, varying resolutions, and multiple spectral bands simultaneously. Current approaches predominantly rely on visual-only modeling paradigms that not only struggle with complex background interference and inherently scarce target features, but also exhibit limited generalization capabilities across complex omni-scene environments where significant domain shifts and appearance variations occur. In this work, we reveal a critical oversight in existing paradigms: the neglect of readily available auxiliary metadata describing imaging parameters and acquisition conditions, such as spectral bands, sensor platforms, resolution, and observation perspectives. To address this limitation, we propose the Auxiliary Metadata Driven Infrared Small Target Detector (AuxDet), a novel multimodal framework that is the first to incorporate metadata into the IRSTD paradigm for scene-aware optimization. Through a high-dimensional fusion module based on multi-layer perceptrons (MLPs), AuxDet dynamically integrates metadata semantics with visual features, guiding adaptive representation learning for each individual sample. Additionally, we design a lightweight prior-initialized enhancement module using 1D convolutional blocks to further refine fused features and recover fine-grained target cues. Extensive experiments on the challenging WideIRSTD-Full benchmark demonstrate that AuxDet consistently outperforms state-of-the-art methods, validating the critical role of auxiliary information in improving robustness and accuracy in omni-domain IRSTD tasks. Code is available at https://github.com/GrokCV/AuxDet.

cs.CV

MooER: LLM-based Speech Recognition and Translation Models from Moore Threads

In this paper, we present MooER, a LLM-based large-scale automatic speech recognition (ASR) / automatic speech translation (AST) model of Moore Threads. A 5000h pseudo labeled dataset containing open source and self collected speech data is used for training. We achieve performance comparable to other open source models trained with up to hundreds of thousands of hours of labeled speech data. Meanwhile, experiments conducted on Covost2 Zh2en testset suggest that our model outperforms other open source Speech LLMs. A BLEU score of 25.2 can be obtained. The main contributions of this paper are summarized as follows. First, this paper presents a training strategy for encoders and LLMs on speech related tasks (including ASR and AST) using a small size of pseudo labeled data without any extra manual annotation and selection. Second, we release our ASR and AST models and plan to open-source our training code and strategy in the near future. Moreover, a model trained on 8wh scale training data is planned to be released later on.

cs.CL

Background Semantics Matter: Cross-Task Feature Exchange Network for Clustered Infrared Small Target Detection

Infrared small target detection presents significant challenges due to the limited intrinsic features of the target and the overwhelming presence of visually similar background distractors. We contend that background semantics are critical for distinguishing between objects that appear visually similar in this context. To address this challenge, we propose a task, clustered infrared small target detection, and introduce DenseSIRST, a benchmark dataset that provides per-pixel semantic annotations for background regions. This dataset facilitates the shift from sparse to dense target detection. This dataset facilitates the shift from sparse to dense target detection. Building on this resource, we propose the Background-Aware Feature Exchange Network (BAFE-Net), a multi-task architecture that jointly tackles target detection and background semantic segmentation. BAFE-Net incorporates a dynamic cross-task feature hard-exchange mechanism, enabling the effective exchange of target and background semantics between the two tasks. Comprehensive experiments demonstrate that BAFE-Net significantly enhances target detection accuracy while mitigating false alarms. The DenseSIRST dataset, along with the code and trained models, is publicly available at https://github.com/GrokCV/BAFE-Net.

cs.CV

Long-time dynamics of stochastic wave equation with dissipative damping and its full discretization: exponential ergodicity and strong law of large numbers

For stochastic wave equation, when the dissipative damping is a non-globally Lipschitz function of the velocity, there are few results on the long-time dynamics, in particular, the exponential ergodicity and strong law of large numbers, for the equation and its numerical discretization to our knowledge. Focus on this issue, the main contributions of this paper are as follows. First, based on constructing novel Lyapunov functionals, we show the unique invariant measure and exponential ergodicity of the underlying equation and its full discretization. Second, the error estimates of invariant measures both in Wasserstein distance and in the weak sense are obtained. Third, the strong laws of large numbers of the equation and the full discretization are obtained, which states that the time averages of the exact and numerical solutions are shown to converge to the ergodic limit almost surely.

math.PR

Building a digital twin of EDFA: a grey-box modeling approach

To enable intelligent and self-driving optical networks, high-accuracy physical layer models are required. The dynamic wavelength-dependent gain effects of non-constant-pump erbium-doped fiber amplifiers (EDFAs) remain a crucial problem in terms of modeling, as it determines optical-to-signal noise ratio as well as the magnitude of fiber nonlinearities. Black-box data-driven models have been widely studied, but it requires a large size of data for training and suffers from poor generalizability. In this paper, we derive the gain spectra of EDFAs as a simple univariable linear function, and then based on it we propose a grey-box EDFA gain modeling scheme. Experimental results show that for both automatic gain control (AGC) and automatic power control (APC) EDFAs, our model built with 8 data samples can achieve better performance than the neural network (NN) based model built with 900 data samples, which means the required data size for modeling can be reduced by at least two orders of magnitude. Moreover, in the experiment the proposed model demonstrates superior generalizability to unseen scenarios since it is based on the underlying physics of EDFAs. The results indicate that building a customized digital twin of each EDFA in optical networks become feasible, which is essential especially for next generation multi-band network operations.

eess.SP

Strong convergence rates for a full discretization of stochastic wave equation with nonlinear damping

The paper establishes the strong convergence rates of a spatio-temporal full discretization of the stochastic wave equation with nonlinear damping in dimension one and two. We discretize the SPDE by applying a spectral Galerkin method in space and a modified implicit exponential Euler scheme in time. The presence of the super-linearly growing damping in the underlying model brings challenges into the error analysis. To address these difficulties, we first achieve upper mean-square error bounds, and then obtain mean-square convergence rates of the considered numerical solution. This is done without requiring the moment bounds of the full approximations. The main result shows that, in dimension one, the scheme admits a convergence rate of order $\tfrac12$ in space and order $1$ in time. In dimension two, the error analysis is more subtle and can be done at the expense of an order reduction due to an infinitesimal factor. Numerical experiments are performed and confirm our theoretical findings.

math.NA

Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire

In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to segment the audio and then transcribe each segmentation. These two stages are addressed separately by speaker change detection (SCD) and automatic speech recognition (ASR). Most previous SCD systems rely solely on speaker information and ignore the importance of speech content. In this paper, we propose a novel SCD system that considers both cues of speaker difference and speech content. These two cues are converted into token-level representations by the continuous integrate-and-fire (CIF) mechanism and then combined for detecting speaker changes on the token acoustic boundaries. We evaluate the performance of our approach on a public real-recorded meeting dataset, AISHELL-4. The experiment results show that our method outperforms a competitive frame-level baseline system by 2.45% equal coverage-purity (ECP). In addition, we demonstrate the importance of speech content and speaker difference to the SCD task, and the advantages of conducting SCD on the token acoustic boundaries compared with conducting SCD frame by frame.

cs.SD

Sequence-level Speaker Change Detection with Difference-based Continuous Integrate-and-fire

Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels.

cs.SD

Strong convergence rates of a fully discrete scheme for the Cahn-Hilliard-Cook equation

The first aim of this paper is to examine existence, uniqueness and regularity for the Cahn-Hilliard-Cook (CHC) equation in space dimension $d\leq 3$. By applying a spectral Galerkin method to the infinite dimensional equation, we elaborate the well-posedness and regularity of the finite dimensional approximate problem. The key idea lies in transforming the stochastic problem {\color{black}{with additive noise}} into an equivalent random equation. The regularity of the solution to the equivalent random equation is obtained, in one dimension, with the aid of the Gagliardo-Nirenberg inequality and done in two and three dimensions, by the energy argument. Further, the approximate solution is shown to be strongly convergent to the unique mild solution of the original CHC equation, whose spatio-temporal regularity can be attained by similar arguments. In addition, a fully discrete approximation of such problem is investigated, performed by the spectral Galerkin method in space and the backward Euler method in time. The previously obtained regularity results of the problem help us to identify strong convergence rates of the fully discrete scheme.

math.NA

Strong convergence rates of an explicit scheme for stochastic Cahn--Hilliard equation with additive noise

In this paper, we propose and analyze an explicit time-stepping scheme for a spatial discretization of stochastic Cahn--Hilliard equation with additive noise. The fully discrete approximation combines a spectral Galerkin method in space with a tamed exponential Euler method in time. In contrast to implicit schemes in the literature, the explicit scheme here is easily implementable and produces significant improvement in the computational efficiency. It is shown that the fully discrete approximation converges strongly to the exact solution, with strong convergence rates identified. Different from the tamed time-stepping schemes for stochastic Allen--Cahn equations, essential difficulties arise in the analysis due to the presence of the unbounded linear operator in front of the nonlinearity. To overcome them, new and non-trivial arguments are developed in the present work. To the best of our knowledge, it is the first result concerning an explicit scheme for the stochastic Cahn--Hilliard equation. Numerical experiments are finally performed to confirm the theoretical results.

math.NA

Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection

Nowadays, most methods in end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may encounter confusion between similar context-specific phrases, which hurts predictions at the token level. In this work, we focus on mitigating confusion problems with fine-grained contextual knowledge selection (FineCoS). In FineCoS, we introduce fine-grained knowledge to reduce the uncertainty of token predictions. Specifically, we first apply phrase selection to narrow the range of phrase candidates, and then conduct token attention on the tokens in the selected phrase candidates. Moreover, we re-normalize the attention weights of most relevant phrases in inference to obtain more focused phrase-level contextual representations, and inject position information to better discriminate phrases or tokens. On LibriSpeech and an in-house 160,000-hour dataset, we explore the proposed methods based on a controllable all-neural biasing method, collaborative decoding (ColDec). The proposed methods provide at most 6.1% relative word error rate reduction on LibriSpeech and 16.4% relative character error rate reduction on the in-house dataset over ColDec.

cs.CL

Weak convergence of the backward Euler method for stochastic Cahn--Hilliard equation with additive noise

We prove a weak rate of convergence of a fully discrete scheme for stochastic Cahn--Hilliard equation with additive noise, where the spectral Galerkin method is used in space and the backward Euler method is used in time. Compared with the Allen--Cahn type stochastic partial differential equation, the error analysis here is much more sophisticated due to the presence of the unbounded operator in front of the nonlinear term. To address such issues, a novel and direct approach has been exploited which does not rely on a Kolmogorov equation but on the integration by parts formula from Malliavin calculus. To the best of our knowledge, the rates of weak convergence are revealed in the stochastic Cahn--Hilliard equation setting for the first time.

math.NA

Improving Pseudo-label Training For End-to-end Speech Recognition Using Gradient Mask

In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose a novel approach to combine their ideas for end-to-end speech recognition model. Without any extra loss function, we utilize the Gradient Mask to optimize the model when training on pseudo-label. This method forces the speech recognition model to predict from the masked input to learn strong acoustic representation and make training robust to label noise. In our semi-supervised experiments, the method can improve the model performance when training on pseudo-label and our method achieved competitive results comparing with other semi-supervised approaches on the Librispeech 100 hours experiments.

eess.AS

Weak convergence rates for an explicit full-discretization of stochastic Allen-Cahn equation with additive noise

We discretize the stochastic Allen-Cahn equation with additive noise by means of a spectral Galerkin method in space and a tamed version of the exponential Euler method in time. The resulting error bounds are analyzed for the spatio-temporal full discretization in both strong and weak senses. Different from existing works, we develop a new and direct approach for the weak error analysis, which does not rely on the use of the associated Kolmogorov equation or Itô's formula and is therefore non-Markovian in nature. Such an approach thus has a potential to be applied to non-Markovian equations such as stochastic Volterra equations or other types of fractional SPDEs, which suffer from the lack of Kolmogorov equations. It turns out that the obtained weak convergence rates are, in both spatial and temporal direction, essentially twice as high as the strong convergence rates. Also, it is revealed how the weak convergence rates depend on the regularity of the noise. Numerical experiments are finally reported to confirm the theoretical conclusion.

math.NA

A Data-Fusion-Assisted Telemetry Layer for Autonomous Optical Networks

For further improving the capacity and reliability of optical networks, a closed-loop autonomous architecture is preferred. Considering a large number of optical components in an optical network and many digital signal processing modules in each optical transceiver, massive real-time data can be collected. However, for a traditional monitoring structure, collecting, storing and processing a large size of data are challenging tasks. Moreover, strong correlations and similarities between data from different sources and regions are not properly considered, which may limit function extension and accuracy improvement. To address abovementioned issues, a data-fusion-assisted telemetry layer between the physical layer and control layer is proposed in this paper. The data fusion methodologies are elaborated on three different levels: Source Level, Space Level and Model Level. For each level, various data fusion algorithms are introduced and relevant works are reviewed. In addition, proof-of-concept use cases for each level are provided through simulations, where the benefits of the data-fusion-assisted telemetry layer are shown.

eess.SP