SearcharxivSearch

arXiv subjects

Anton Schick

Publications and source records attributed to Anton Schick.

9 recordsLinked to original sources

Theoretical Properties of Multivariate Random Forest in Feature Selection and its Application to Facial Morphology-Gene Detection

This work establishes a theoretical foundation for joint feature selection with multivariate outcomes, positioning the permutation-based variable importance measure (PVIM) of multivariate random forests (MRF) as a principled tool for high-dimensional feature selection. We establish the first consistency guaranty for MRF, showing that it retains all truly influential features with probability tending to one as the sample size grows to infinity under mild regularity conditions. Incomplete U-statistics is employed to incorporate three layers of randomness: subsampling of subjects for training each tree, subsampling of features at each split, and permutation of each feature for the out-of-bag (OOB) samples. Unlike independence-based screening that evaluates each feature in isolation, PVIM is a joint screening approach that accounts for multicollinearity, nonlinear, high-order interactions, and subject heterogeneity via ensemble aggregation. Moreover, we demonstrate the practical utility of MRF through a genome-wide association study (GWAS) of human facial morphology (with 2,342 subjects and 453,273 SNPs), where MRF identifies several novel loci and interaction hubs that extend prior findings. Extensive simulations show that MRF accurately identifies truly influential signals while producing parsimonious feature sets with well controlled false selection rates, outperforming canonical correlation analysis (CCA) and several other independence multivariate screening approaches. In addition, we also propose a novel simulation framework, including image outcomes, that more closely mimic the intricate nature of real-world data and provide rigorous testbeds for machine learning research.

stat.ME

Estimating the error distribution function in nonparametric regression

We construct an efficient estimator for the error distribution function of the nonparametric regression model Y = r(Z) + e. Our estimator is a kernel smoothed empirical distribution function based on residuals from an under-smoothed local quadratic smoother for the regression function.

math.ST

Uniform convergence of convolution estimators for the response density in nonparametric regression

We consider a nonparametric regression model $Y=r(X)+\varepsilon$ with a random covariate $X$ that is independent of the error $\varepsilon$. Then the density of the response $Y$ is a convolution of the densities of $\varepsilon$ and $r(X)$. It can therefore be estimated by a convolution of kernel estimators for these two densities, or more generally by a local von Mises statistic. If the regression function has a nowhere vanishing derivative, then the convolution estimator converges at a parametric rate. We show that the convergence holds uniformly, and that the corresponding process obeys a functional central limit theorem in the space $C_0(\mathbb {R})$ of continuous functions vanishing at infinity, endowed with the sup-norm. The estimator is not efficient. We construct an additive correction that makes it efficient.

math.ST

Empirical likelihood approach to goodness of fit testing

Motivated by applications to goodness of fit testing, the empirical likelihood approach is generalized to allow for the number of constraints to grow with the sample size and for the constraints to use estimated criteria functions. The latter is needed to deal with nuisance parameters. The proposed empirical likelihood based goodness of fit tests are asymptotically distribution free. For univariate observations, tests for a specified distribution, for a distribution of parametric form, and for a symmetric distribution are presented. For bivariate observations, tests for independence are developed.

math.ST

The transfer principle: A tool for complete case analysis

This paper gives a general method for deriving limiting distributions of complete case statistics for missing data models from corresponding results for the model where all data are observed. This provides a convenient tool for obtaining the asymptotic behavior of complete case versions of established full data methods without lengthy proofs. The methodology is illustrated by analyzing three inference procedures for partially linear regression models with responses missing at random. We first show that complete case versions of asymptotically efficient estimators of the slope parameter for the full model are efficient, thereby solving the problem of constructing efficient estimators of the slope parameter for this model. Second, we derive an asymptotically distribution free test for fitting a normal distribution to the errors. Finally, we obtain an asymptotically distribution free test for linearity, that is, for testing that the nonparametric component of these models is a constant. This test is new both when data are fully observed and when data are missing at random.

math.ST

Optimality of estimators for misspecified semi-Markov models

Suppose we observe a geometrically ergodic semi-Markov process and have a parametric model for the transition distribution of the embedded Markov chain, for the conditional distribution of the inter-arrival times, or for both. The first two models for the process are semiparametric, and the parameters can be estimated by conditional maximum likelihood estimators. The third model for the process is parametric, and the parameter can be estimated by an unconditional maximum likelihood estimator. We determine heuristically the asymptotic distributions of these estimators and show that they are asymptotically efficient. If the parametric models are not correct, the (conditional) maximum likelihood estimators estimate the parameter that maximizes the Kullback--Leibler information. We show that they remain asymptotically efficient in a nonparametric sense.

math.ST

Uniformly root-$N$ consistent density estimators for weakly dependent invertible linear processes

Convergence rates of kernel density estimators for stationary time series are well studied. For invertible linear processes, we construct a new density estimator that converges, in the supremum norm, at the better, parametric, rate $n^{-1/2}$. Our estimator is a convolution of two different residual-based kernel estimators. We obtain in particular convergence rates for such residual-based kernel estimators; these results are of independent interest.

math.ST

Efficient prediction for linear and nonlinear autoregressive models

Conditional expectations given past observations in stationary time series are usually estimated directly by kernel estimators, or by plugging in kernel estimators for transition densities. We show that, for linear and nonlinear autoregressive models driven by independent innovations, appropriate smoothed and weighted von Mises statistics of residuals estimate conditional expectations at better parametric rates and are asymptotically efficient. The proof is based on a uniform stochastic expansion for smoothed and weighted von Mises processes of residuals. We consider, in particular, estimation of conditional distribution functions and of conditional quantile functions.

math.ST

Estimating invariant laws of linear processes by U-statistics

Suppose we observe an invertible linear process with independent mean-zero innovations and with coefficients depending on a finite-dimensional parameter, and we want to estimate the expectation of some function under the stationary distribution of the process. The usual estimator would be the empirical estimator. It can be improved using the fact that the innovations are centered. We construct an even better estimator using the representation of the observations as infinite-order moving averages of the innovations. Then the expectation of the function under the stationary distribution can be written as the expectation under the distribution of an infinite series in terms of the innovations, and it can be estimated by a U-statistic of increasing order (also called an ``infinite-order U-statistic'') in terms of the estimated innovations. The estimator can be further improved using the fact that the innovations are centered. This improved estimator is optimal if the coefficients of the linear process are estimated optimally.

math.ST