Searcharxiv⌕ Search

arXiv subjects

Ka Long Keith Ho

Publications and source records attributed to Ka Long Keith Ho.

5 recordsLinked to original sources

Approximating Simple ReLU Networks based on Spectral Decomposition of Fisher Information

Properties of Fisher information matrices of 2-layer neural ReLU networks with random hidden weights are studied. For these networks, it is known that the eigenvalue distribution highly concentrates on several eigenspaces approximately. In particular, the eigenvalues for the first three eigenspaces account for 97.7% of the trace of the Fisher information matrix, independently of the number of parameters. In this paper, we identify the function spaces which correspond to those major eigenspaces. This function space consists of the spherical harmonic functions whose orders are not greater than 2. This result relates to the Mercer decomposition of the neural tangent kernels.

stat.ML↗

Neural Tangent Kernels and Fisher Information Matrices for Simple ReLU Networks with Random Hidden Weights

Fisher information matrices and neural tangent kernels (NTK) for 2-layer ReLU networks with random hidden weight are argued. We discuss the relation between both notions as a linear transformation and show that spectral decomposition of NTK with concrete forms of eigenfunctions with major eigenvalues. We also obtain an approximation formula of the functions presented by the 2-layer neural networks.

cs.LG↗

Variable selection via thresholding

Variable selection comprises an important step in many modern statistical inference procedures. In the regression setting, when estimators cannot shrink irrelevant signals to zero, covariates without relationships to the response often manifest small but non-zero regression coefficients. The ad hoc procedure of discarding variables whose coefficients are smaller than some threshold is often employed in practice. We formally analyze a version of such thresholding procedures and develop a simple thresholding method that consistently estimates the set of relevant variables under mild regularity assumptions. Using this thresholding procedure, we propose a sparse, $\sqrt{n}$-consistent and asymptotically normal estimator whose non-zero elements do not exhibit shrinkage. The performance and applicability of our approach are examined via numerical studies of simulated and real data.

math.ST↗

Adaptive Ridge Approach to Heteroscedastic Regression

We propose an adaptive ridge (AR) estimation scheme for a heteroscedastic linear regression model with log-linear noise in data. We simultaneously estimate the mean and variance parameters, demonstrating new asymptotic distributional and tightness properties in a sparse setting. We also show that estimates for zero parameters shrink with more iterations under suitable assumptions for tuning parameters. Aspects of application and possible generalizations are presented through simulations and real data examples.

math.ST↗

Small Area Estimation under Square Root Transformed Fay-Herriot model with Functional Measurement Error in Covariates

We consider a small area estimation model under square-root transformation in the presence of functional measurement error. When measurement error is present, the Bayes predictor can no longer be used as it depends on the covariates even if parameters are known. Therefore suitable replacements are called for, and we propose a predictor that only depends on observed responses and data obtained from a large secondary survey. Moreover, some estimating methods of unknown parameters are considered. In the simulations section, We evaluate the performance using the mean squared prediction error (MSPE) and discuss several scenarios in terms of the number of areas and the sample size in a large secondary survey.

stat.ME↗