Searcharxiv⌕ Search

arXiv subjects

Guanjie Lyu

Publications and source records attributed to Guanjie Lyu.

9 recordsLinked to original sources

Bernstein-smoothed estimation and bootstrap inference for the lower-tail Spearman's rho curve

This paper studies a Bernstein-smoothed plug-in estimator for the lower-tail Spearman's rho curve, a rank-based measure of local concordance defined through a normalized copula integral over the lower-left square $[0,p]^2$. The estimator applies the lower-tail Spearman functional to the empirical Bernstein copula and introduces a degree parameter that controls finite-sample regularization. The contribution is target-specific estimation and inference for the lower-tail Spearman's rho curve rather than a new general-purpose copula estimator. We establish uniform strong consistency on compact intervals away from zero whenever $m \to \infty$. Under a target-specific integrated Bernstein-bias condition and $\sqrt n/m \to 0$, we derive functional weak convergence with the same first-order Gaussian limit as the empirical copula-based estimator; this degree regime includes $m = \lfloor n^{2/3}\rfloor$. For inference, we justify a smoothed beta bootstrap based on sampling from the empirical beta copula and construct pointwise confidence intervals from the absolute centered bootstrap root. Monte Carlo experiments for the Farlie--Gumbel--Morgenstern, Gaussian, Clayton, and Frank copulas assess pointwise coverage over the threshold grid and show that Bernstein smoothing generally reduces integrated variance and often lowers mean integrated squared error under weak to moderate dependence. Additional common-sample experiments compare the proposed estimator with the empirical beta and empirical checkerboard Bernstein copula estimators in terms of pointwise error, integrated error, computation time, selected degrees, numerical stability, and pointwise coverage. A sensitivity analysis shows that $m = \lfloor n^{2/3}\rfloor$ is a simple theoretically admissible default that avoids the most severe oversmoothing. A descriptive application to the Loss--ALAE insurance claims data illustrates our method.

math.ST↗

Dirichlet kernel density estimation on the simplex with missing data

Nonparametric density estimation for compositional data supported on the simplex is examined under a missing at random mechanism. Rather than imputing missing values and estimating the density from a completed data set, we adopt a strategy based on inverse probability weighting. The proposed estimator uses an adaptive Dirichlet kernel, which ensures nonnegativity on the simplex and favorable behavior near the boundary. When the observation probabilities are unknown, they are estimated through a Nadaraya-Watson regression step. The large-sample properties of the estimator are derived, including pointwise bias and variance expansions, optimal smoothing rates, and asymptotic normality. A simulation study investigates its finite-sample performance under varying sample sizes and missing rates. Simulations show our method outperforms inverse-probability-weighted kernel density estimators based on additive and isometric log-ratio transformations of the data for certain target densities. The methodology is further illustrated through an application to leukocyte composition data from the National Health and Nutrition Examination Survey (NHANES), which allows for the identification of the modal immune profile in the sampled population.

stat.ME↗

Asymptotic properties of the multivariate Szász-Mirakyan estimator for cumulative distribution functions on the nonnegative orthant

The asymptotic properties of multivariate Szász-Mirakyan estimators for cumulative distribution functions (cdf) supported on the nonnegative orthant are investigated. Explicit bias and variance expansions are derived on compact subsets of the interior, yielding sharp mean squared error characterizations and optimal smoothing rates. The analysis shows that the proposed Poisson smoothing yields a non-negligible variance reduction relative to the empirical cdf, leading to asymptotic efficiency gains that can be quantified through local and global deficiency measures. The behavior of the estimator near the boundary of its support is examined separately. Under a boundary-layer scaling that preserves nondegenerate Poisson smoothing as the evaluation point approaches the boundary of $[0,\infty)^d$, bias and variance expansions are obtained that differ fundamentally from those in the interior region. In particular, the variance reduction mechanism disappears at leading order, implying that no asymptotically optimal smoothing parameter exists in the boundary regime. Central limit theorems and almost sure uniform consistency are also established. Together, these results provide a unified asymptotic theory for multivariate Szász-Mirakyan cdf estimation and clarify the distinct roles of smoothing in the interior and boundary regions.

math.ST↗

Tweedie-based nonparametric estimation for semicontinuous mixed densities

Semicontinuous outcomes occur frequently in health services, insurance, and cost studies. Standard nonparametric density estimators are not well suited to such data because they do not naturally accommodate the mixed structure, the nonnegative support, or the pronounced boundary effects near zero. To address these limitations, we introduce an asymmetric kernel estimator for mixed densities on $[0,\infty)$ based on the Tweedie distribution. For a power parameter $p\in(1,2)$, the Tweedie kernel itself has a point mass at zero and an absolutely continuous component on $(0,\infty)$, yielding a unified smoothing construction that preserves the atom at zero and smooths the positive component using the full semicontinuous sample. We establish pointwise bias and variance expansions, derive asymptotic formulae for the mean squared error and mean integrated squared error, obtain optimal bandwidth rates, and prove asymptotic normality. We propose a profile least-squares cross-validation procedure to jointly select the bandwidth and the power parameter. Simulation results show competitive performance, particularly in challenging boundary-spike and heavy-tailed settings, and an application to emergency department length-of-stay data illustrates the practical value of the method.

stat.ME↗

Semiparametric copula-based quantile regression for semicontinuous outcomes with application to healthcare data

A semiparametric copula-based two-part quantile regression framework is developed for the analysis of semicontinuous outcomes characterized by a point mass at zero and a continuous positive component. The proposed approach models the occurrence and magnitude processes separately and links them through copula-based conditional distributions, allowing for flexible dependence structures and nonlinear covariate effects across quantiles. Large-sample properties of the resulting estimator are established, and extensive simulation studies demonstrate improved finite-sample performance relative to logistic/linear quantile regression, particularly under nonlinear dependence and substantial zero inflation. An application to healthcare data illustrates how the proposed method provides a nuanced characterization of the association between social deprivation and uncompensated and charity care burdens, revealing heterogeneous and nonlinear effects that are not captured by competing approaches.

stat.ME↗

Rank-based concordance for zero-inflated data: New representations, estimators, and sharp bounds

Quantifying concordance between two random variables is crucial in applications. Traditional estimation techniques for commonly used concordance measures, such as Gini's gamma or Spearman's rho, often fail when data contain ties. This is particularly problematic for zero-inflated data, characterized by a combination of discrete mass in zero and a continuous component, which frequently appear in insurance, weather forecasting, and biomedical applications. This study provides a new formulation of Gini's gamma and Spearman's footrule, two rank-based concordance measures that incorporate absolute rank differences, tailored to zero-inflated continuous distributions. Along the way, we correct an expression of Spearman's rho for zero-inflated data previously presented in the literature. The best-possible upper and lower bounds for these measures in zero-inflated continuous settings are established, making the estimators useful and interpretable in practice. We pair our theoretical results with simulations and two real-life applications in insurance and weather forecasting, respectively. Our results illustrate the impact of zero inflation on dependence estimation, emphasizing the benefits of appropriately adjusted zero-inflated measures.

stat.ME↗

Testing equality between two-sample dependence structure using Bernstein polynomials

Tests of equality of copulas between two samples are introduced and studied using the empirical Bernstein copula process. Three statistics are proposed and their asymptotic properties are established. Besides, a subsampling Bernstein version method is investigated and compared with multiplier bootstrap. Simulation study showed that the Bernstein tests outperform the tests based on the empirical copula.

math.ST↗

Testing Symmetry for Bivariate Copulas using Bernstein Polynomials

In this work, tests of symmetry for bivariate copulas are introduced and studied using empirical Bernstein copula process. Three statistics are proposed and their asymptotic properties are established. Besides, a multiplier bootstrap Bernstein version is investigated for implementation purpose. The simulation study demonstrated the superior performance of the Bernstein tests compared to tests based on empirical copulas. Furthermore, in real data applications, these tests consistently yielded similar conclusions across a diverse range of scenarios.

stat.ME↗