SearcharxivSearch

arXiv subjects

Akito Narahara

Publications and source records attributed to Akito Narahara.

2 recordsLinked to original sources

Bootstrap Error Estimation and Sketch-Size Selection for Sketched Ridge Regression

Randomized sketching reduces the computational cost of large ridge-regression problems, but the coefficient error depends on the realized sketch. We extend the paired-row bootstrap from randomized least squares to ridge regression, enabling coefficient-error estimation using only compressed data. Under a fixed coefficient dimension and increasing data and sketch sizes, we derive asymptotic linear representations and Gaussian limits for the sketched estimator and the conditional bootstrap distribution. These results establish uniform consistency of the bootstrap error distribution and asymptotically exact coverage when the estimator and error bound are computed from the same sketch. For sketches with independent, mean-zero, variance-one entries, an explicit covariance formula separates the effects of the residual, regularization, and fourth moment of the sketch entries, and shows that Rademacher entries minimize the leading covariance matrix in the Loewner order. We also develop a fast linearized bootstrap, an order-statistic correction for finitely many bootstrap replicates, and a Bonferroni rule for selecting from a fixed set of sketch sizes. Experiments on two real and two synthetic data sets support the proposed methods. With a sketch size 15 times the number of coefficients, 199 bootstrap replicates, and nominal coverage of 95\%, the bootstrap with refitting attains coverage between 92.0\% and 95.3\%; after the order-statistic correction, coverage ranges from 94.7\% to 97.7\%.

math.ST

Mixture Proportion Estimation and Weakly-supervised Kernel Test for Conditional Independence

Mixture proportion estimation (MPE) aims to estimate class priors from unlabeled data. This task is a critical component in weakly supervised learning, such as PU learning, learning with label noise, and domain adaptation. Existing MPE methods rely on the \textit{irreducibility} assumption or its variant for identifiability. In this paper, we propose novel assumptions based on conditional independence (CI) given the class label, which ensure identifiability even when irreducibility does not hold. We develop method of moments estimators under these assumptions and analyze their asymptotic properties. Furthermore, we present weakly-supervised kernel tests to validate the CI assumptions, which are of independent interest in applications such as causal discovery and fairness evaluation. Empirically, we demonstrate the improved performance of our estimators compared with existing methods and that our tests successfully control both type I and type II errors.\label{key}

cs.LG