SearcharxivSearch

arXiv subjects

Chunlin Wang

Publications and source records attributed to Chunlin Wang.

At least 19 recordsLinked to original sources

A semiparametric two-sample homogeneity test with nonignorable nonresponse using callback data

Testing the homogeneity of two distributions is fundamental in statistics, but classical procedures may fail under nonignorable nonresponse. In many surveys, callback data record repeated contact attempts and provide auxiliary information about the response mechanism. We develop a semiparametric framework for two-sample homogeneity testing that explicitly incorporates such information. The response mechanism is modeled by a flexible semiparametric callback model, while the two population distributions are linked through a density ratio model. Within this unified framework, we propose an empirical likelihood ratio test for distributional homogeneity and show that, under the null hypothesis, it has a Wilks-type chi-square limit. To facilitate computation, we develop an efficient expectation-maximization-type algorithm. Simulation results show that the proposed method controls type I error well and achieves substantially higher power than existing methods that ignore nonignorable missingness. An application to real survey income data illustrates its practical value.

stat.ME

Semiparametric Joint Inference for Sensitivity and Specificity at the Youden-Optimal Cut-Off

Sensitivity and specificity evaluated at an optimal diagnostic cut-off are fundamental measures of classification accuracy when continuous biomarkers are used for disease diagnosis. Joint inference for these quantities is challenging because their estimators are evaluated at a common, data-driven threshold estimated from both diseased and healthy samples, inducing statistical dependence. Existing approaches are largely based on parametric assumptions or fully nonparametric procedures, which may be sensitive to model misspecification or lack efficiency in moderate samples. We propose a semiparametric framework for joint inference on sensitivity and specificity at the Youden-optimal cut-off under the density ratio model. Using maximum empirical likelihood, we derive estimators of the optimal threshold and the corresponding sensitivity and specificity, and establish their joint asymptotic normality. This leads to Wald-type and range-preserving logit-transformed confidence regions. Simulation studies show that the proposed method achieves accurate coverage with improved efficiency relative to existing parametric and nonparametric alternatives across a variety of distributional settings. An analysis of COVID-19 antibody data demonstrates the practical advantages of the proposed approach for diagnostic decision-making.

stat.ME

Semiparametric inference for inequality measures under nonignorable nonresponse using callback data

This paper develops semiparametric methods for estimation and inference of widely used inequality measures when survey data are subject to nonignorable nonresponse, a challenging setting in which response probabilities depend on the unobserved outcomes. Such nonresponse mechanisms are common in household surveys and invalidate standard inference procedures due to selection bias and lack of population representativeness. We address this problem by exploiting callback data from repeated contact attempts and adopting a semiparametric model that leaves the outcome distribution unspecified. We construct semiparametric full-likelihood estimators for the underlying distribution and the associated inequality measures, and establish their large-sample properties for a broad class of functionals, including quantiles, the Theil index, and the Gini index. Explicit asymptotic variance expressions are derived, enabling valid Wald-type inference under nonignorable nonresponse. To facilitate implementation, we propose a stable and computationally convenient expectation-maximization algorithm, whose steps either admit closed-form expressions or reduce to fitting a standard logistic regression model. Simulation studies demonstrate that the proposed procedures effectively correct nonresponse bias and achieve near-benchmark efficiency. An application to Consumer Expenditure Survey data illustrates the practical gains from incorporating callback information when making inference on inequality measures.

econ.EM

Some ergodic theorems over squarefree numbers and squarefull numbers

In 2022, Bergelson and Richter gave a new dynamical generalization of the prime number theorem by establishing an ergodic theorem along the number of prime factors of integers. They also showed that this generalization holds as well if the integers are restricted to be squarefree. In this paper, we present the concept of invariant averages under multiplications for arithmetic functions. Utilizing the properties of these invariant averages, we derive several ergodic theorems over squarefree numbers and squarefull numbers. These theorems have significant connections to the Erd\H{o}s-Kac Theorem, the Bergelson-Richter Theorem, and the Loyd Theorem.

math.NT

Expansions in completions of global function fields

It is well known that any power series over a finite field represents a rational function if and only if its sequence of coefficients is ultimately periodic. The famous Christol's Theorem states that a power series over a finite field is algebraic if and only if its sequence of coefficients is $p$-automatic. In this paper, we extend these two results to expansions of elements in the completion of a global function field under a nontrivial valuation. As application of our generalization of Christol's theorem, we answer some questions about $\beta$-expansions of formal Laurent series over finite fields.

math.NT

On Erd\H{o}s covering systems in global function fields

A covering system of the integers is a finite collection of arithmetic progressions whose union is the set of integers. A well-known problem on covering systems is the minimum modulus problem posed by Erd\H{o}s in 1950, who asked whether the minimum modulus in such systems with distinct moduli can be arbitrarily large. This problem was resolved by Hough in 2015, who showed that the minimum modulus is at most $10^{16}$. In 2022, Balister, Bollob\'as, Morris, Sahasrabudhe and Tiba reduced Hough's bound to $616,000$ by developing Hough's method. They call it the distortion method. In this paper, by applying this method, we mainly prove that there does not exist any covering system of multiplicity $s$ in any global function field of genus $g$ over $\mathbb{F}_q$ for $q\geq (1.14+0.16g)e^{6.5+0.97g}s^2$. In particular, there is no covering system of $\mathbb{F}_q[x]$ with distinct moduli for $q\geq 759$.

math.NT

On covering systems of polynomial rings over finite fields

In 1950, Erd\H{o}s posed a question known as the minimum modulus problem on covering systems for $\mathbb{Z}$, which asked whether the minimum modulus of a covering system with distinct moduli is bounded. This long-standing problem was finally resolved by Hough in 2015, as he proved that the minimum modulus of any covering system with distinct moduli does not exceed $10^{16}$. Recently, Balister, Bollob\'as, Morris, Sahasrabudhe, and Tiba developed a versatile method called the distortion method and significantly reduced Hough's bound to $616,000$. In this paper, we apply this method to present a proof that the smallest degree of the moduli in any covering system for $\mathbb{F}_q[x]$ of multiplicity $s$ is bounded by a constant depending only on $s$ and $q$. Consequently, we successfully resolve the minimum modulus problem for $\mathbb{F}_q[x]$ and disprove a conjecture by Azlin.

math.NT

Prompt Guided Transformer for Multi-Task Dense Prediction

Task-conditional architecture offers advantage in parameter efficiency but falls short in performance compared to state-of-the-art multi-decoder methods. How to trade off performance and model parameters is an important and difficult problem. In this paper, we introduce a simple and lightweight task-conditional model called Prompt Guided Transformer (PGT) to optimize this challenge. Our approach designs a Prompt-conditioned Transformer block, which incorporates task-specific prompts in the self-attention mechanism to achieve global dependency modeling and parameter-efficient feature adaptation across multiple tasks. This block is integrated into both the shared encoder and decoder, enhancing the capture of intra- and inter-task features. Moreover, we design a lightweight decoder to further reduce parameter usage, which accounts for only 2.7% of the total model parameters. Extensive experiments on two multi-task dense prediction benchmarks, PASCAL-Context and NYUD-v2, demonstrate that our approach achieves state-of-the-art results among task-conditional methods while using fewer parameters, and maintains a significant balance between performance and parameter size.

cs.CV

Hypothesis test on a mixture forward-incubation-time epidemic model with application to COVID-19 outbreak

The distribution of the incubation period of the novel coronavirus disease that emerged in 2019 (COVID-19) has crucial clinical implications for understanding this disease and devising effective disease-control measures. Qin et al. (2020) designed a cross-sectional and forward follow-up study to collect the duration times between a specific observation time and the onset of COVID-19 symptoms for a number of individuals. They further proposed a mixture forward-incubation-time epidemic model, which is a mixture of an incubation-period distribution and a forward time distribution, to model the collected duration times and to estimate the incubation-period distribution of COVID-19. In this paper, we provide sufficient conditions for the identifiability of the unknown parameters in the mixture forward-incubation-time epidemic model when the incubation period follows a two-parameter distribution. Under the same setup, we propose a likelihood ratio test (LRT) for testing the null hypothesis that the mixture forward-incubation-time epidemic model is a homogeneous exponential distribution. The testing problem is non-regular because a nuisance parameter is present only under the alternative. We establish the limiting distribution of the LRT and identify an explicit representation for it. The limiting distribution of the LRT under a sequence of local alternatives is also obtained. Our simulation results indicate that the LRT has desirable type I errors and powers, and we analyze a COVID-19 outbreak dataset from China to illustrate the usefulness of the LRT.

stat.ME

Distribution of residues of an algebraic number modulo ideals of degree one

Let $f(x)$ be an irreducible polynomial with integer coefficients of degree at least two. Hooley proved that the roots of the congruence equation $f(x)\equiv 0\mod n$ is uniformly distributed. as a parallel of Hooley's theorem under ideal theoretical setting, we prove the uniformity of the distribution of residues of an algebraic number modulo degree one ideals. Then using this result we show that the roots of a system of polynomial congruences are uniformly distributed. Finally, the distribution of digits of n-adic expansions of an algebraic number is discussed.

math.NT

Newton polygons for $L$-functions of generalized Kloosterman sums

In this paper, we study the Newton polygons for the $L$-functions of $n$-variable generalized Kloosterman sums. Generally, the Newton polygon has a topological lower bound, called the Hodge polygon. In order to determine the Hodge polygon, we explicitly construct a basis of the top dimensional Dwork cohomology. Using Wan's decomposition theorem and diagonal local theory, we obtain when the Newton polygon coincides with the Hodge polygon. In particular, we concretely get the slope sequence for $L$-function of $\bar{F}(\barλ,x):=\sum_{i=1}^nx_i^{a_i}+\barλ\prod_{i=1}^nx_i^{-1}$.

math.NT

Semiparametric inference on general functionals of two semicontinuous populations

In this paper, we propose new semiparametric procedures for making inference on linear functionals and their functions of two semicontinuous populations. The distribution of each population is usually characterized by a mixture of a discrete point mass at zero and a continuous skewed positive component, and hence such distribution is semicontinuous in the nature. To utilize the information from both populations, we model the positive components of the two mixture distributions via a semiparametric density ratio model. Under this model setup, we construct the maximum empirical likelihood estimators of the linear functionals and their functions, and establish the asymptotic normality of the proposed estimators. We show the proposed estimators of the linear functionals are more efficient than the fully nonparametric ones. The developed asymptotic results enable us to construct confidence regions and perform hypothesis tests for the linear functionals and their functions. We further apply these results to several important summary quantities such as the moments, the mean ratio, the coefficient of variation, and the generalized entropy class of inequality measures. Simulation studies demonstrate the advantages of our proposed semiparametric method over some existing methods. Two real data examples are provided for illustration.

stat.ME

An efficient explicit full-discrete scheme for strong approximation of stochastic Allen-Cahn equation

In Becker and Jentzen (2019) and Becker et al. (2017), an explicit temporal semi-discretization scheme and a space-time full-discretization scheme were, respectively, introduced and analyzed for the additive noise-driven stochastic Allen-Cahn type equations, with strong convergence rates recovered. The present work aims to propose a different explicit full-discrete scheme to numerically solve the stochastic Allen-Cahn equation with cubic nonlinearity, perturbed by additive space-time white noise. The approximation is easily implementable, performing the spatial discretization by a spectral Galerkin method and the temporal discretization by a kind of nonlinearity-tamed accelerated exponential integrator scheme. Error bounds in a strong sense are analyzed for both the spatial semi-discretization and the spatio-temporal full discretization, with convergence rates in both space and time explicitly identified. It turns out that the obtained convergence rate of the new scheme is, in the temporal direction, twice as high as existing ones in the literature. Numerical results are finally reported to confirm the previous theoretical findings.

math.NA

A novel semi-supervised multi-view clustering framework for screening Parkinson's disease

In recent years, there are many research cases for the diagnosis of Parkinson's disease (PD) with the brain magnetic resonance imaging (MRI) by utilizing the traditional unsupervised machine learning methods and the supervised deep learning models. However, unsupervised learning methods are not good at extracting accurate features among MRIs and it is difficult to collect enough data in the field of PD to satisfy the need of training deep learning models. Moreover, most of the existing studies are based on single-view MRI data, of which data characteristics are not sufficient enough. In this paper, therefore, in order to tackle the drawbacks mentioned above, we propose a novel semi-supervised learning framework called Semi-supervised Multi-view learning Clustering architecture technology (SMC). The model firstly introduces the sliding window method to grasp different features, and then uses the dimensionality reduction algorithms of Linear Discriminant Analysis (LDA) to process the data with different features. Finally, the traditional single-view clustering and multi-view clustering methods are employed on multiple feature views to obtain the results. Experiments show that our proposed method is superior to the state-of-art unsupervised learning models on the clustering effect. As a result, it may be noted that, our work could contribute to improving the effectiveness of identifying PD by previous labeled and subsequent unlabeled medical MRI data in the realistic medical environment.

eess.IV

Asymptotic coverage probabilities of bootstrap percentile confidence intervals for constrained parameters

The asymptotic behaviour of the commonly used bootstrap percentile confidence interval is investigated when the parameters are subject to linear inequality constraints. We concentrate on the important one- and two-sample problems with data generated from general parametric distributions in the natural exponential family. The focus of this paper is on quantifying the coverage probabilities of the parametric bootstrap percentile confidence intervals, in particular their limiting behaviour near boundaries. We propose a local asymptotic framework to study this subtle coverage behaviour. Under this framework, we discover that when the true parameters are on, or close to, the restriction boundary, the asymptotic coverage probabilities can always exceed the nominal level in the one-sample case; however, they can be, remarkably, both under and over the nominal level in the two-sample case. Using illustrative examples, we show that the results provide theoretical justification and guidance on applying the bootstrap percentile method to constrained inference problems.

math.ST

Criterion for the integrality of hypergeometric series with parameters from quadratic fields

For the hypergeometric series with parameters from the rational fields, there is an effective criterion due to Christol to decide whether the hypergeometric series is N-integral or not. Christol criterion is a basic and vital tool in the recent striking work of Delaygue, Rivoal and Roques on the N-integrality of the hypergeometric mirror maps with rational parameters. In this paper, we develop a systematic theory on the N-integrality of the hypergeometric series with parameters from quadratic fields. We first present a detailed $p$-adic analysis to set up a criterion of the $p$-adic integrality of the hypergeometric series with parameters from rational fields. Consequently, we present two equivalent statements for the hypergeometric series with parameters from algebraic number fields to be N-integral. Finally, by using these results, introducing a new function that extends the Christol's function and developing a further $p$-adic analysis, we establish a criterion for the N-integrality of the hypergeometric series with parameters from the quadratic fields. In the process, there are two important ingredients. One is the uniform distribution result of roots of a quadratic congruence which is due to Duke, Friedlander and Iwaniec together with Toth. Another one is an upper bound on the number of solutions of polynomial congruences given by Stewart in 1991.

math.NT

The elementary symmetric functions of reciprocals of the elements of arithmetic progressions

Let $a$ and $b$ be positive integers. In 1946, Erdős and Niven proved that there are only finitely many positive integers $n$ for which one or more of the elementary symmetric functions of $1/b, 1/(a+b),..., 1/(an-a+b)$ are integers. In this paper, we show that for any integer $k$ with $1\le k\le n$, the $k$-th elementary symmetric function of $1/b, 1/(a+b),..., 1/(an-a+b)$ is not an integer except that either $b=n=k=1$ and $a\ge 1$, or $a=b=1, n=3$ and $k=2$. This refines the Erdős-Niven theorem and answers an open problem raised by Chen and Tang in 2012.

math.NT

The elementary symmetric functions of a reciprocal polynomial sequence

Erdös and Niven proved in 1946 that for any positive integers $m$ and $d$, there are at most finitely many integers $n$ for which at least one of the elementary symmetric functions of $1/m, 1/(m+d), ..., 1/(m+(n-1)d)$ are integers. Recently, Wang and Hong refined this result by showing that if $n\geq 4$, then none of the elementary symmetric functions of $1/m, 1/(m+d), ..., 1/(m+(n-1)d)$ is an integer for any positive integers $m$ and $d$. Let $f$ be a polynomial of degree at least $2$ and of nonnegative integer coefficients. In this paper, we show that none of the elementary symmetric functions of $1/f(1), 1/f(2), ..., 1/f(n)$ is an integer except for $f(x)=x^{m}$ with $m\geq2$ being an integer and $n=1$.

math.NT