SearcharxivSearch

arXiv subjects

Adam Krzyżak

Publications and source records attributed to Adam Krzyżak.

2 recordsLinked to original sources

Estimation of a regression function from dependent data by over-parametrized deep neural networks learned by gradient descent

Estimation of a regression function from exponentially $\beta$-mixing data is considered. The $L_2$ error with integration with respect to the design is used as the error criterion. Deep neural network estimates with logistic activation function are defined, where all parameters are learned by gradient descent. The rate of convergence of the expected $L_2$ error is analyzed for $(p,C)$-smooth regression functions. In the special case that the design is concentrated on a $d^*$-dimensional manifold, it is shown that the expected $L_2$ error of the estimate achieves a rate of convergence which depends on $d^*$ and not on the dimension $d$ of the design.

math.ST

Learning of deep neural network regression estimates using gradient descent with pruning

Estimation of a regression function from independent and identically distributed data is considered. The $L_2$ error with integration with respect to the design variable is used as the error criterion. An initially randomly pruned fully connected deep neural network with logistic squasher as activation function is fitted to the data via gradient descent, using a data-dependent choice of non-constant stepsizes during gradient descent. It is shown that this network achieves (up to a logarithmic factor) the optimal minimax rate of convergence in case that the regression function is $(p,C)$--smooth. Here the estimate is able to circumvent the curse of dimensionality provided the predictors are concentrated in the neighborhood of a low dimensional manifold. The finite sample size performance of the estimate is illustrated by applying it to the simulated data.

math.ST