SearcharxivSearch

arXiv subjects

Niklas Dexheimer

Publications and source records attributed to Niklas Dexheimer.

9 recordsLinked to original sources

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model

A refined statistical understanding of LLM pre-training requires the analysis of the transformer architecture for data distributions that encapsulate key characteristics of text data. To address this, we propose a text data distribution based on an extension of the log-bilinear language model from the natural language processing literature. For this data generating process, we derive generalization bounds for deep transformer architectures, highlighting the dependence on the network architecture, the vocabulary size, the number of documents and the document length.

math.ST

Sparse Estimation for High-Dimensional L\'evy-driven Ornstein--Uhlenbeck Processes from Discrete Observations

We study high-dimensional drift estimation for L\'evy-driven Ornstein--Uhlenbeck processes based on discrete observations. Assuming sparsity of the drift matrix, we analyze Lasso and Slope estimators constructed from approximate likelihoods and derive sharp nonasymptotic oracle inequalities. Our bounds disentangle the contributions of discretization error and stochastic fluctuations, and establish minimax optimal convergence rates under suitable choices of tuning parameters in a high-frequency regime. We further quantify the sample complexity required to attain these rates depending on the L\'evy noise. The results extend the theory of high-dimensional statistics for stochastic processes to a substantially broader class of noise mechanisms, in particular pure jump processes. They also demonstrate that Lasso and Slope remain competitive for jump-driven systems, providing practical guidance for inference in applications where L\'evy processes are a natural modeling choice.

math.ST

Spike-timing-dependent Hebbian learning as noisy gradient descent

Hebbian learning is a key principle underlying learning in biological neural networks. We relate a Hebbian spike-timing-dependent plasticity rule to noisy gradient descent with respect to a non-convex loss function on the probability simplex. Despite the constant injection of noise and the non-convexity of the underlying optimization problem, one can rigorously prove that the considered Hebbian learning dynamic identifies the presynaptic neuron with the highest activity and that the convergence is exponentially fast in the number of iterations. This is non-standard and surprising as typically noisy gradient descent with fixed noise level only converges to a stationary regime where the noise causes the dynamic to fluctuate around a minimiser.

cs.LG

Improving the Convergence Rates of Forward Gradient Descent with Repeated Sampling

Forward gradient descent (FGD) has been proposed as a biologically more plausible alternative of gradient descent as it can be computed without backward pass. Considering the linear model with $d$ parameters, previous work has found that the prediction error of FGD is, however, by a factor $d$ slower than the prediction error of stochastic gradient descent (SGD). In this paper we show that by computing $\ell$ FGD steps based on each training sample, this suboptimality factor becomes $d/(\ell \wedge d)$ and thus the suboptimality of the rate disappears if $\ell \gtrsim d.$ We also show that FGD with repeated sampling can adapt to low-dimensional structure in the input distribution. The main mathematical challenge lies in controlling the dependencies arising from the repeated sampling process.

math.ST

Data-driven optimal stopping: A pure exploration analysis

The standard theory of optimal stopping is based on the idealised assumption that the underlying process is essentially known. In this paper, we drop this restriction and study data-driven optimal stopping for a general diffusion process, focusing on investigating the statistical performance of the proposed estimator of the optimal stopping barrier. More specifically, we derive non-asymptotic upper bounds on the simple regret, along with uniform and non-asymptotic PAC bounds. Minimax optimality is verified by completing the upper bound results with matching lower bounds on the simple regret. All results are shown both under general conditions on the payoff functions and under more refined assumptions that mimic the margin condition used in binary classification, leading to an improved rate of convergence. Additionally, we investigate how our results on the simple regret transfer to the cumulative regret for a specific exploration-exploitation strategy, both with respect to lower bounds and upper bounds.

math.ST

Adaptive nonparametric drift estimation for multivariate jump diffusions under sup-norm risk

We investigate nonparametric drift estimation for multidimensional jump diffusions based on continuous observations. The results are derived under anisotropic smoothness assumptions and the estimators' performance is measured in terms of the sup-norm loss. We present two different Nadaraya--Watson type estimators, which are both shown to achieve the classical nonparametric rate of convergence under varying assumptions on the jump measure. Fully data-driven versions of both estimators are also introduced and shown to attain the same rate of convergence. The results rely on novel uniform moment bounds for empirical processes associated to the investigated jump diffusion, which are of independent interest.

math.ST

Estimating the characteristics of stochastic damping Hamiltonian systems from continuous observations

We consider nonparametric invariant density and drift estimation for a class of multidimensional degenerate resp. hypoelliptic diffusion processes, so-called stochastic damping Hamiltonian systems or kinetic diffusions, under anisotropic smoothness assumptions on the unknown functions. The analysis is based on continuous observations of the process, and the estimators' performance is measured in terms of the sup-norm loss. Regarding invariant density estimation, we obtain highly nonclassical results for the rate of convergence, which reflect the inhomogeneous variance structure of the process. Concerning estimation of the drift vector, we suggest both non-adaptive and fully data-driven procedures. All of the aforementioned results strongly rely on tight uniform moment bounds for empirical processes associated to deterministic and stochastic integrals of the investigated process, which are also proven in this paper.

math.ST

On Lasso and Slope drift estimators for Lévy-driven Ornstein--Uhlenbeck processes

We investigate the problem of estimating the drift parameter of a high-dimensional Lévy-driven Ornstein--Uhlenbeck process under sparsity constraints. It is shown that both Lasso and Slope estimators achieve the minimax optimal rate of convergence (up to numerical constants), for tuning parameters chosen independently of the confidence level, which improves the previously obtained results for standard Ornstein--Uhlenbeck processes. The results are nonasymptotic and hold both in probability and conditional expectation with respect to an event resembling the restricted eigenvalue condition.

math.ST

Mixing it up: A general framework for Markovian statistics

Up to now, the nonparametric analysis of multidimensional continuous-time Markov processes has focussed strongly on specific model choices, mostly related to symmetry of the semigroup. While this approach allows to study the performance of estimators for the characteristics of the process in the minimax sense, it restricts the applicability of results to a rather constrained set of stochastic processes and in particular hardly allows incorporating jump structures. As a consequence, for many models of applied and theoretical interest, no statement can be made about the robustness of typical statistical procedures beyond the beautiful, but limited framework available in the literature. To close this gap, we identify $β$-mixing of the process and heat kernel bounds on the transition density as a suitable combination to obtain $\sup$-norm and $L^2$ kernel invariant density estimation rates matching the case of reversible multidimenisonal diffusion processes and outperforming density estimation based on discrete i.i.d. or weakly dependent data. Moreover, we demonstrate how up to $\log$-terms, optimal $\sup$-norm adaptive invariant density estimation can be achieved within our general framework based on tight uniform moment bounds and deviation inequalities for empirical processes associated to additive functionals of Markov processes. The underlying assumptions are verifiable with classical tools from stability theory of continuous time Markov processes and PDE techniques, which opens the door to evaluate statistical performance for a vast amount of Markov models. We highlight this point by showing how multidimensional jump SDEs with Lévy driven jump part under different coefficient assumptions can be seamlessly integrated into our framework, thus establishing novel adaptive $\sup$-norm estimation rates for this class of processes.

math.ST