SearcharxivSearch

arXiv subjects

Takayuki Kawashima

Publications and source records attributed to Takayuki Kawashima.

8 recordsLinked to original sources

Bootstrap Error Estimation and Sketch-Size Selection for Sketched Ridge Regression

Randomized sketching reduces the computational cost of large ridge-regression problems, but the coefficient error depends on the realized sketch. We extend the paired-row bootstrap from randomized least squares to ridge regression, enabling coefficient-error estimation using only compressed data. Under a fixed coefficient dimension and increasing data and sketch sizes, we derive asymptotic linear representations and Gaussian limits for the sketched estimator and the conditional bootstrap distribution. These results establish uniform consistency of the bootstrap error distribution and asymptotically exact coverage when the estimator and error bound are computed from the same sketch. For sketches with independent, mean-zero, variance-one entries, an explicit covariance formula separates the effects of the residual, regularization, and fourth moment of the sketch entries, and shows that Rademacher entries minimize the leading covariance matrix in the Loewner order. We also develop a fast linearized bootstrap, an order-statistic correction for finitely many bootstrap replicates, and a Bonferroni rule for selecting from a fixed set of sketch sizes. Experiments on two real and two synthetic data sets support the proposed methods. With a sketch size 15 times the number of coefficients, 199 bootstrap replicates, and nominal coverage of 95\%, the bootstrap with refitting attains coverage between 92.0\% and 95.3\%; after the order-statistic correction, coverage ranges from 94.7\% to 97.7\%.

math.ST

Convergence fragility in probit Bayesian kernel machine regression implemented in the bkmr R package for binary-outcome environmental mixture analyses: a simulation study

Background. Bayesian kernel machine regression (BKMR) is widely used for exposure-mixture analyses with binary outcomes through a probit extension. Because a bkmr fit can complete without providing adequate effective posterior information, simulation studies should separate execution success from MCMC convergence diagnostics. Methods. We evaluated the public bkmr probit workflow using bkmr::SimData() for data generation, bkmr::kmbayes() for model fitting, and posterior for convergence diagnostics. The balanced generator used family = "binomial", hfun = 2, beta.true = 0.5, ind = 1:2, and M = 4. SimData() generated the covariate as X = 3*cos(z1) + 2*rnorm(n). Four chains were initialized with chain-specific randomized starting values generated reproducibly from the fixed initial-value base seed 20260621. These values affected only the initial state of the sampler and did not alter the BKMR model, default priors, or Metropolis-Hastings proposals. Results. Of 431 prespecified tasks, 430 returned fitted objects and one task had a numerical non-completion. Diagnostic adequacy was limited: rank-normalized R-hat <= 1.01 threshold was achieved in 55/431 tasks, bulk-ESS >= 400 in 85/431, tail-ESS >= 400 in 44/431, and both ESS criteria in 44/431. The primary diagnostic criterion, R-hat at or below the 1.01 threshold with both bulk-ESS and tail-ESS >= 400, was met in 30/431 prespecified tasks, corresponding to 30/430 completed fits. Conclusions. Completion of probit BKMR fits in bkmr should not be equated with convergence of the retained MCMC draws. Applied analyses should report the number of chains, warmup and retained iterations, rank-normalized R-hat, bulk-ESS, and tail-ESS rather than rely on a fixed iteration count or on fit completion alone.

stat.AP

New Insight of Spatial Scan Statistics via Regression Model

The spatial scan statistic is widely used to detect disease clusters in epidemiological surveillance. Since the seminal work by~\cite{kulldorff1997}, numerous extensions have emerged, including methods for defining scan regions, detecting multiple clusters, and expanding statistical models. Notably,~\cite{jung2009} and~\cite{ZHANG20092851} introduced a regression-based approach accounting for covariates, encompassing classical methods such as those of~\cite{kulldorff1997}. Another key extension is the expectation-based approach~\citep{neill2005anomalous,neillphdthesis}, which differs from the population-based approach represented by~\cite{kulldorff1997} in terms of hypothesis testing. In this paper, we bridge the regression-based approach with both expectation-based and population-based approaches. We reveal that the two approaches are separated by a simple difference: the presence or absence of an intercept term in the regression model. Exploiting the above simple difference, we propose new spatial scan statistics under the Gaussian and Bernoulli models. We further extend the regression-based approach by incorporating the well-known sparse L0 penalty and show that the derivation of spatial scan statistics can be expressed as an equivalent optimization problem. Our extended framework accommodates extensions such as space-time scan statistics and detecting multiple clusters while naturally connecting with existing spatial regression-based cluster detection. Considering the relation to case-specific models~\citep{she2011,10.1214/11-STS377}, clusters detected by spatial scan statistics can be viewed as outliers in terms of robust statistics. Numerical experiments with real data illustrate the behavior of our proposed statistics under specified settings.

stat.ME

Robust Estimation for Kernel Exponential Families with Smoothed Total Variation Distances

In statistical inference, we commonly assume that samples are independent and identically distributed from a probability distribution included in a pre-specified statistical model. However, such an assumption is often violated in practice. Even an unexpected extreme sample called an {\it outlier} can significantly impact classical estimators. Robust statistics studies how to construct reliable statistical methods that efficiently work even when the ideal assumption is violated. Recently, some works revealed that robust estimators such as Tukey's median are well approximated by the generative adversarial net (GAN), a popular learning method for complex generative models using neural networks. GAN is regarded as a learning method using integral probability metrics (IPM), which is a discrepancy measure for probability distributions. In most theoretical analyses of Tukey's median and its GAN-based approximation, however, the Gaussian or elliptical distribution is assumed as the statistical model. In this paper, we explore the application of GAN-like estimators to a general class of statistical models. As the statistical model, we consider the kernel exponential family that includes both finite and infinite-dimensional models. To construct a robust estimator, we propose the smoothed total variation (STV) distance as a class of IPMs. Then, we theoretically investigate the robustness properties of the STV-based estimators. Our analysis reveals that the STV-based estimator is robust against the distribution contamination for the kernel exponential family. Furthermore, we analyze the prediction accuracy of a Monte Carlo approximation method, which circumvents the computational difficulty of the normalization constant.

stat.ML

Stochastic Gradient Descent for Stochastic Doubly-Nonconvex Composite Optimization

The stochastic gradient descent has been widely used for solving composite optimization problems in big data analyses. Many algorithms and convergence properties have been developed. The composite functions were convex primarily and gradually nonconvex composite functions have been adopted to obtain more desirable properties. The convergence properties have been investigated, but only when either of composite functions is nonconvex. There is no convergence property when both composite functions are nonconvex, which is named the \textit{doubly-nonconvex} case.To overcome this difficulty, we assume a simple and weak condition that the penalty function is \textit{quasiconvex} and then we obtain convergence properties for the stochastic doubly-nonconvex composite optimization problem.The convergence rate obtained here is of the same order as the existing work.We deeply analyze the convergence rate with the constant step size and mini-batch size and give the optimal convergence rate with appropriate sizes, which is superior to the existing work. Experimental results illustrate that our method is superior to existing methods.

stat.ML

On Difference Between Two Types of $γ$-divergence for Regression

The $γ$-divergence is well-known for having strong robustness against heavy contamination. By virtue of this property, many applications via the $γ$-divergence have been proposed. There are two types of \gd\ for regression problem, in which the treatments of base measure are different. In this paper, we compare them and pointed out a distinct difference between these two divergences under heterogeneous contamination where the outlier ratio depends on the explanatory variable. One divergence has the strong robustness under heterogeneous contamination. The other does not have in general, but has when the parametric model of the response variable belongs to a location-scale family in which the scale does not depend on the explanatory variables or under homogeneous contamination where the outlier ratio does not depend on the explanatory variable. \citet{hung.etal.2017} discussed the strong robustness in a logistic regression model with an additional assumption that the tuning parameter $γ$ is sufficiently large. The results obtained in this paper hold for any parametric model without such an additional assumption.

math.ST

Robust and Sparse Regression in GLM by Stochastic Optimization

The generalized linear model (GLM) plays a key role in regression analyses. In high-dimensional data, the sparse GLM has been used but it is not robust against outliers. Recently, the robust methods have been proposed for the specific example of the sparse GLM. Among them, we focus on the robust and sparse linear regression based on the $γ$-divergence. The estimator of the $γ$-divergence has strong robustness under heavy contamination. In this paper, we extend the robust and sparse linear regression based on the $γ$-divergence to the robust and sparse GLM based on the $γ$-divergence with a stochastic optimization approach in order to obtain the estimate. We adopt the randomized stochastic projected gradient descent as a stochastic optimization approach and extend the established convergence property to the classical first-order necessary condition. By virtue of the stochastic optimization approach, we can efficiently estimate parameters for very large problems. Particularly, we show the linear regression, logistic regression and Poisson regression with $L_1$ regularization in detail as specific examples of robust and sparse GLM. In numerical experiments and real data analysis, the proposed method outperformed comparative methods.

stat.ML

Robust and Sparse Regression via $γ$-divergence

In high-dimensional data, many sparse regression methods have been proposed. However, they may not be robust against outliers. Recently, the use of density power weight has been studied for robust parameter estimation and the corresponding divergences have been discussed. One of such divergences is the $γ$-divergence and the robust estimator using the $γ$-divergence is known for having a strong robustness. In this paper, we consider the robust and sparse regression based on $γ$-divergence. We extend the $γ$-divergence to the regression problem and show that it has a strong robustness under heavy contamination even when outliers are heterogeneous. The loss function is constructed by an empirical estimate of the $γ$-divergence with sparse regularization and the parameter estimate is defined as the minimizer of the loss function. To obtain the robust and sparse estimate, we propose an efficient update algorithm which has a monotone decreasing property of the loss function. Particularly, we discuss a linear regression problem with $L_1$ regularization in detail. In numerical experiments and real data analyses, we see that the proposed method outperforms past robust and sparse methods.

stat.ME