SearcharxivSearch

arXiv subjects

Qiying Wang

Publications and source records attributed to Qiying Wang.

14 recordsLinked to original sources

Limit theorems for a class of martingale arrays with applications in nonlinear cointegrating regression

This paper develops a new asymptotic theory for a broad class of martingales, establishing convergence to limiting distributions that involve a functional of stochastic integrals. The proposed limit theorem substantially extends existing martingale asymptotic theory by accommodating a wider class of dependence structures. As a primary application, the theory is applied to nonlinear regression models with nonstationary time series, yielding a rigorous framework for asymptotic inference on nonlinear least square estimators.

math.ST

ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has emerged as a transformative workload in cloud clusters, enabling large language models (LLMs) to solve complex problems through interactions with real world. However, unlike traditional RL, agentic RL demands substantial external cloud resources, e.g., CPUs for code execution and GPUs for reward models, that exist outside the primary training cluster. Existing agentic RL framework typically rely on static over-provisioning, i.e., resources are often tied to long-lived trajectories or isolated by tasks, which leads to severe resource inefficiency. We propose the action-level orchestration, and incorporate it into ARL-Tangram, a unified resource management system that enables fine-grained external resource sharing and elasticity. ARL-Tangram utilizes a unified action-level formulation and an elastic scheduling algorithm to minimize action completion time (ACT) while satisfying heterogeneous resource constraints. Further, heterogeneous resource managers are tailored to efficiently support the action-level execution on resources with heterogeneous characteristics and topologies. Evaluation on real-world agentic RL tasks demonstrates that ARL-Tangram improves average ACT by up to 4.3$\times$, speeds up the step duration of RL training by up to 1.5$\times$, and saves the external resources by up to 71.2$\%$. This system has been deployed to support the training of the MiMo series models.

cs.DC

MiMo-V2-Flash Technical Report

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that interleaves Sliding Window Attention (SWA) with global attention, with a 128-token sliding window under a 5:1 hybrid ratio. The model is pre-trained on 27 trillion tokens with Multi-Token Prediction (MTP), employing a native 32k context length and subsequently extended to 256k. To efficiently scale post-training compute, MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm. In this framework, domain-specialized teachers (e.g., trained via large-scale reinforcement learning) provide dense and token-level reward, enabling the student model to perfectly master teacher expertise. MiMo-V2-Flash rivals top-tier open-weight models such as DeepSeek-V3.2 and Kimi-K2, despite using only 1/2 and 1/3 of their total parameters, respectively. During inference, by repurposing MTP as a draft model for speculative decoding, MiMo-V2-Flash achieves up to 3.6 acceptance length and 2.6x decoding speedup with three MTP layers. We open-source both the model weights and the three-layer MTP weights to foster open research and community collaboration.

cs.CL

MiMo-Audio: Audio Language Models are Few-Shot Learners

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown that scaling next-token prediction pretraining enables strong generalization capabilities in text, and we believe this paradigm is equally applicable to the audio domain. By scaling MiMo-Audio's pretraining data to over one hundred million of hours, we observe the emergence of few-shot learning capabilities across a diverse set of audio tasks. We develop a systematic evaluation of these capabilities and find that MiMo-Audio-7B-Base achieves SOTA performance on both speech intelligence and audio understanding benchmarks among open-source models. Beyond standard metrics, MiMo-Audio-7B-Base generalizes to tasks absent from its training data, such as voice conversion, style transfer, and speech editing. MiMo-Audio-7B-Base also demonstrates powerful speech continuation capabilities, capable of generating highly realistic talk shows, recitations, livestreaming and debates. At the post-training stage, we curate a diverse instruction-tuning corpus and introduce thinking mechanisms into both audio understanding and generation. MiMo-Audio-7B-Instruct achieves open-source SOTA on audio understanding benchmarks (MMSU, MMAU, MMAR, MMAU-Pro), spoken dialogue benchmarks (Big Bench Audio, MultiChallenge Audio) and instruct-TTS evaluations, approaching or surpassing closed-source models. Model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-Audio.

cs.CL

Locally trimmed least squares: conventional inference in possibly nonstationary models

A novel IV estimation method, that we term Locally Trimmed LS (LTLS), is developed which yields estimators with (mixed) Gaussian limit distributions in situations where the data may be weakly or strongly persistent. In particular, we allow for nonlinear predictive type of regressions where the regressor can be stationary short/long memory as well as nonstationary long memory process or a nearly integrated array. The resultant t-tests have conventional limit distributions (i.e. N(0; 1)) free of (near to unity and long memory) nuisance parameters. In the case where the regressor is a fractional process, no preliminary estimator for the memory parameter is required. Therefore, the practitioner can conduct inference while being agnostic about the exact dependence structure in the data. The LTLS estimator is obtained by applying certain chronological trimming to the OLS instrument via the utilisation of appropriate kernel functions of time trend variables. The finite sample performance of LTLS based t-tests is investigated with the aid of a simulation experiment. An empirical application to the predictability of stock returns is also provided.

econ.EM

Uniform convergence rates for a class of martingales with application in non-linear cointegrating regression

For a class of martingales, this paper provides a framework on the uniform consistency with broad applicability. The main condition imposed is only related to the conditional variance of the martingale, which holds true for stationary mixing time series, stationary iterated random function, Harris recurrent Markov chains and $I(1)$ processes with innovations being a linear process. Using the established results, this paper investigates the uniform convergence of the Nadaraya-Watson estimator in a non-linear cointegrating regression model. Our results not only provide sharp convergence rate, but also the optimal range for the uniform convergence to be held. This paper also considers the uniform upper and lower bound estimates for a functional of Harris recurrent Markov chain, which are of independent interests.

math.ST

Long-range dependent time series specification

In this paper we propose using a nonparametric model specification test for parametric time series with long-range dependence (LRD). To establish asymptotic distributions of the proposed test statistic, we develop new central limit theorems for certain weighted quadratic forms of stationary time series with LRD. To implement our proposed test in practice, we develop a computer-intensive parametric bootstrap simulation procedure for finding simulated critical values. As a result, our finite-sample studies demonstrate that both the proposed theory and the simulation procedure work well, and that the proposed test has little size distortion and reasonable power.

math.ST

Self-normalized Cramér type moderate deviations for the maximum of sums

Let $X_1,X_2,...$ be independent random variables with zero means and finite variances, and let $S_n=\sum_{i=1}^nX_i$ and $V^2_n=\sum_{i=1}^nX^2_i$. A Cramér type moderate deviation for the maximum of the self-normalized sums $\max_{1\leq k\leq n}S_k/V_n$ is obtained. In particular, for identically distributed $X_1,X_2,...,$ it is proved that $P(\max_{1\leq k\leq n}S_k\geq xV_n)/(1-Φ(x))\rightarrow2$ uniformly for $0<x\leq\mathrm{o}(n^{1/6})$ under the optimal finite third moment of $X_1$.

math.ST

A specification test for nonlinear nonstationary models

We provide a limit theory for a general class of kernel smoothed U-statistics that may be used for specification testing in time series regression with nonstationary data. The test framework allows for linear and nonlinear models with endogenous regressors that have autoregressive unit roots or near unit roots. The limit theory for the specification test depends on the self-intersection local time of a Gaussian process. A new weak convergence result is developed for certain partial sums of functions involving nonstationary time series that converges to the intersection local time process. This result is of independent interest and is useful in other applications. Simulations examine the finite sample performance of the test.

math.ST

Strong approximations of level exceedences related to multiple hypothesis testing

Particularly in genomics, but also in other fields, it has become commonplace to undertake highly multiple Student's $t$-tests based on relatively small sample sizes. The literature on this topic is continually expanding, but the main approaches used to control the family-wise error rate and false discovery rate are still based on the assumption that the tests are independent. The independence condition is known to be false at the level of the joint distributions of the test statistics, but that does not necessarily mean, for the small significance levels involved in highly multiple hypothesis testing, that the assumption leads to major errors. In this paper, we give conditions under which the assumption of independence is valid. Specifically, we derive a strong approximation that closely links the level exceedences of a dependent ``studentized process'' to those of a process of independent random variables. Via this connection, it can be seen that in high-dimensional, low sample-size cases, provided the sample size diverges faster than the logarithm of the number of tests, the assumption of independent $t$-tests is often justified.

math.ST

On weighted approximations in $D[0, 1]$ with applications to self-normalized partial sum processes

Let $X, X_1, X_2,...$ be a sequence of non-degenerate i.i.d. random variables with mean zero. The best possible weighted approximations are investigated in $D[0, 1]$ for the partial sum processes $\{S_{[nt]}, 0\le t\le 1\}$, where $S_n=\sum_{j=1}^nX_j$, under the assumption that $X$ belongs to the domain of attraction of the normal law. The conclusions then are used to establish similar results for the sequence of self-normalized partial sum processes $\{S_{[nt]}/V_n, 0\le t\le 1\}$, where $V_n^2=\sum_{j=1}^nX_j^2$. $L_p$ approximations of self-normalized partial sum processes are also discussed.

math.PR

Asymptotics of Studentized U-type processes for changepoint problems

This paper investigates weighted approximations for studentized $U$-statistics type processes, both with symmetric and antisymmetric kernels, only under the assumption that the distribution of the projection variate is in the domain of attraction of the normal law. The classical second moment condition $E|h(X_1,X_2)|^2 < \infty$ is also relaxed in both cases. The results can be used for testing the null assumption of having a random sample versus the alternative that there is a change in distribution in the sequence.

math.PR

Cramér-type large deviations for samples from a finite population

Cramér-type large deviations for means of samples from a finite population are established under weak conditions. The results are comparable to results for the so-called self-normalized large deviation for independent random variables. Cramér-type large deviations for the finite population Student $t$-statistic are also investigated.

math.ST

Exact convergence rate and leading term in central limit theorem for student's t statistic

The leading term in the normal approximation to the distribution of Student's t statistic is derived in a general setting, with the sole assumption being that the sampled distribution is in the domain of attraction of a normal law. The form of the leading term is shown to have its origin in the way in which extreme data influence properties of the Studentized sum. The leading-term approximation is used to give the exact rate of convergence in the central limit theorem up to order n^{-1/2}, where n denotes sample size. It is proved that the exact rate uniformly on the whole real line is identical to the exact rate on sets of just three points. Moreover, the exact rate is identical to that for the non-Studentized sum when the latter is normalized for scale using a truncated form of variance, but when the corresponding truncated centering constant is omitted. Examples of characterizations of convergence rates are also given. It is shown that, in some instances, their validity uniformly on the whole real line is equivalent to their validity on just two symmetric points.

math.PR