Estimation of a regression function from dependent data by over-parametrized deep neural networks learned by gradient descent
Estimation of a regression function from exponentially $\beta$-mixing data is considered. The $L_2$ error with integration with respect to the design is used as the error criterion. Deep neural network estimates with logistic activation function are defined, where all parameters are learned by gradient descent. The rate of convergence of the expected $L_2$ error is analyzed for $(p,C)$-smooth regression functions. In the special case that the design is concentrated on a $d^*$-dimensional manifold, it is shown that the expected $L_2$ error of the estimate achieves a rate of convergence which depends on $d^*$ and not on the dimension $d$ of the design.