SearcharxivSearch

arXiv subjects

Benny Avelin

Publications and source records attributed to Benny Avelin.

At least 19 recordsLinked to original sources

Absorption cutoff and stationary singularities for rounded Gaussian random dynamical systems

We study Gaussian random dynamical systems with coordinatewise $\tanh$ nonlinearity, where finite precision is modeled by nearest-grid rounding after each step. Gaussian symmetry reduces the dynamics to an exact Markov chain for the normalized squared radius. Rounding makes the origin absorbing, and the total variation distance to the absorbing equilibrium equals the survival probability of the absorption time. At fixed width, we identify the critical gain and prove an absorption cutoff with Gaussian profile as the mesh tends to zero. At fixed precision, global contraction yields a large-dimension absorption cutoff, while positive drift produces metastability. In the supercritical regime, we prove a large-dimension cutoff to a nonzero invariant law and show that, at fixed dimension, its mass near the repelling origin has a power-law asymptotic. All six main results are formalized in Lean 4 on top of Mathlib and independently checked against restatements that import only Mathlib.

math.PR

On the Effectiveness of Classical Regression Methods for Optimal Switching Problems

Simple regression methods provide robust, near-optimal solutions for optimal switching problems, including high-dimensional ones (up to 50). While the theory requires solving intractable PDE systems, the Longstaff-Schwartz algorithm with classical regression methods achieves excellent switching decisions without extensive hyperparameter tuning. Testing linear models (OLS, Ridge, LASSO), tree-based methods (random forests, gradient boosting), $k$-nearest neighbors, and feedforward neural networks on four benchmark problems, we find that several simple methods maintain stable performance across diverse problem characteristics, outperforming the neural networks we tested against. In our comparison, $k$-NN regression performs consistently well, and with minimal hyperparameter tuning. We establish concentration bounds for this regressor and show that PCA enables $k$-NN to scale to high dimensions.

math.OC

$χ^2$-cut-off phenomenon for Galerkin projections of Fokker-Planck equations with monomial potentials

In this manuscript, we establish the existence/non-existence of the cut-off phenomenon for the Langevin--Kolmogorov random dynamics with monomial convex potentials, possible singular, and driven by a Brownian motion with small strength. We consider a truncated $χ^2$-distance, that is, a distance based on Galerkin projections of the eigensystem, and show that not only a refined knowledge of the eigenvalues is needed but also a refined asymptotics of the growth for the eigenfunctions of the Fokker--Planck equations associated to the Langevin--Kolmogorov dynamics. In addition, this explicit analysis yields asymptotics of the mixing times and, in some regimes, information on the limiting profile, going beyond the product condition and the cut-off window alone.

math.PR

How iteration order influences convergence and stability in deep learning

Despite exceptional achievements, training neural networks remains computationally expensive and is often plagued by instabilities that can degrade convergence. While learning rate schedules can help mitigate these issues, finding optimal schedules is time-consuming and resource-intensive. This work explores theoretical issues concerning training stability in the constant-learning-rate (i.e., without schedule) and small-batch-size regime. Surprisingly, we show that the composition order of gradient updates affects stability and convergence in gradient-based optimizers. We illustrate this new line of thinking using backward-SGD, which produces parameter iterates at each step by reverting the usual forward composition order of batch gradients. Our theoretical analysis shows that in contractive regions (e.g., around minima) backward-SGD converges to a point while the standard forward-SGD generally only converges to a distribution. This leads to improved stability and convergence which we demonstrate experimentally. While full backward-SGD is computationally intensive in practice, it highlights that the extra freedom of modifying the usual iteration composition by reusing creatively previous batches at each optimization step may have important beneficial effects in improving training. Our experiments provide a proof of concept supporting this phenomenon. To our knowledge, this represents a new and unexplored avenue in deep learning optimization.

cs.LG

Exploring Singularities in point clouds with the graph Laplacian: An explicit approach

We develop theory and methods that use the graph Laplacian to analyze the geometry of the underlying manifold of datasets. Our theory provides theoretical guarantees and explicit bounds on the functional forms of the graph Laplacian when it acts on functions defined close to singularities of the underlying manifold. We use these explicit bounds to develop tests for singularities and propose methods that can be used to estimate geometric properties of singularities in the datasets.

stat.ML

Coarse-grained ellipticity and De Giorgi-Nash-Moser theory

We prove local boundedness and a Harnack inequality for nonnegative weak solutions of the equation $-\nabla\cdot(\mathbf{a}(x)\nabla u)=0$ under a coarse-grained ellipticity assumption on the symmetric coefficient field $\mathbf{a}$. Coarse-grained ellipticity is a scale-dependent condition, defined for fields with only $\mathbf{a},\mathbf{a}^{-1}\in L^1$, in terms of families of effective diffusion matrices on triadic cubes of all sizes, and our estimates depend quantitatively on a corresponding coarse-grained ellipticity ratio. We show that coarse-grained ellipticity can be enforced by purely negative Sobolev regularity hypotheses: if $\mathbf{a}\in L^1\cap W^{-s,p}(U)$ and $\mathbf{a}^{-1}\in L^1\cap W^{-t,q}(U)$ for exponents $p,q\in[1,\infty]$ and $s,t\in[0,1)$ satisfying $s<1-\frac{1}{p}$, $t<1-\frac{1}{q}$ and \[ \frac{s+t}{2} + \frac{d}{2}\Bigl(\frac{1}{p}+\frac{1}{q}\Bigr) < 1, \] then $\mathbf{a}$ is coarse-grained elliptic in $U$ and every nonnegative solution satisfies a quantitative unit-scale Harnack inequality. In particular, when $s=t=0$ we recover Trudinger's classical result under the integrability condition $\mathbf{a}\in L^p$, $\mathbf{a}^{-1}\in L^q$ with $\frac{1}{p}+\frac{1}{q}<\frac{2}{d}$, and we obtain the sharp scaling of the Harnack constant in terms of $\|\mathbf{a}\|_{L^p}$ and $\|\mathbf{a}^{-1}\|_{L^q}$. More importantly, our criteria apply to new classes of degenerate and singular coefficient fields for which $\mathbf{a},\mathbf{a}^{-1}\notin L^{1+δ}$ for all $δ>0$, including examples generated by singular fractal measures and Gaussian multiplicative chaos, beyond the reach of previous approaches based solely on integrability assumptions.

math.AP

1D stochastic pressure equation with log-correlated Gaussian coefficients

We study unique solvability for one dimensional stochastic pressure equation with diffusion coefficient given by the Wick exponential of log-correlated Gaussian fields. We prove well-posedness for Dirichlet, Neumann and periodic boundary data, and the initial value problem, covering the cases of both the Wick renormalization of the diffusion and of point-wise multiplication. We provide explicit representations for the solutions in both cases, characterized by the $S$-transform and the Gaussian multiplicative chaos measure.

math.PR

Non-parametric estimation of non-linear diffusion coefficient in parabolic SPDEs

In this article, we introduce a novel non-parametric predictor, based on conditional expectation, for the unknown diffusion coefficient function $σ$ in the stochastic partial differential equation $Lu = σ(u)\dot{W}$, where $L$ is a parabolic second order differential operator and $\dot{W}$ is a suitable Gaussian noise. We prove consistency and derive an upper bound for the error in the $L^p$ norm, in terms of discretization and smoothening parameters $h$ and $\varepsilon$. We illustrate the applicability of the approach and the role of the parameters with several interesting numerical examples.

math.ST

Weak and Perron Solutions for Stationary Kramers-Fokker-Planck Equations in Bounded Domains

In this paper, we investigate weak solutions and Perron-Wiener-Brelot solutions to the linear stationary Kramers-Fokker-Planck equation in bounded domains. We establish the existence of weak solutions in product domains by applying the Lions-Lax-Milgram theorem and the vanishing viscosity method. Furthermore, we show that these solutions coincide in well-behaved domains. Building on the existence of weak solutions in product domains, we develop the foundational theory of Perron-Wiener-Brelot solutions in arbitrary bounded domains. Our results rely on recent advancements in the theory of kinetic Fokker-Planck equations with rough coefficients.

math.AP

Renormalized stochastic pressure equation with log-correlated Gaussian coefficients

We study periodic solutions to the following divergence-form stochastic partial differential equation with Wick-renormalized gradient on the $d$-dimensional flat torus $\mathbb{T}^d$, \[ -\nabla\cdot\left(e^{\diamond (- βX) }\diamond\nabla U\right)=\nabla \cdot (e^{\diamond (- βX)} \diamond \mathbf{F}), \] where $X$ is the log-correlated Gaussian field, $\mathbf{F}$ is a random vector field representing the flux, the in/out-flow of fluid per unit area per unit time, and $\diamond$ denotes the Wick product. The problem is a variant of the stochastic pressure equation, in which $U$ is modeling the pressure of a creeping water-flow in crustal rock that occurs in enhanced geothermal heating. In the original model, the Wick exponential term $e^{\diamond(-βX)}$ is modeling the random permeability of the rock. The porosity field is given by a log-correlated Gaussian random field $βX$, where $β<\sqrt{d}$. We use elliptic regularity theory in order to define a notion of a solution to this (a priori very ill-posed) problem, via modifying the $S$-transform from Gaussian white noise analysis, and then establish the existence and uniqueness of solutions. Moreover, we show that the solution to the problem can be expressed in terms of the Gaussian multiplicative chaos measure.

math.PR

On the Impact of Approximation Errors on Extreme Quantile Estimation with Applications to Functional Data Analysis

We study the effect of approximation errors in assessing the extreme behavior of heavy-tailed random objects. We give conditions for the approximation error such that the standard asymptotic results hold for the classical Hill estimator and the corresponding extreme quantile estimator. As an application, we consider the effect of discretization errors in the computation of the $L^p$-norms related to functional data. We approximate the norms both with Riemann sums and with Monte Carlo integration. We quantify connections between the number of observed functions, the number of discretization points, and the regularity of the underlying functions. In addition, we derive a new concentration inequality for order statistics. This, to the best of our knowledge, is the first Chernoff-type concentration inequality for order statistics presented in the literature that provides an explicit rate at which the ratio between order statistics and tail quantile function converges to one. In our application, the bound is used to provide concentration inequalities measuring the distance between the Hill estimator based on approximated norms and the one based on the true ones.

math.ST

A note on the capacity estimate in metastability for generic configurations

In this paper we further develop the ideas from Geometric Function Theory initially introduced in [arXiv:2206.13206], to derive capacity estimate in metastability for arbitrary configurations. The novelty of this paper is twofold. First, the graph theoretical connection enables us to exactly compute the pre-factor in the capacity. Second, we complete the method from [arXiv:2206.13206] by providing an upper bound using Geometric Function Theory together with Thompson's principle, avoiding explicit constructions of test functions.

math.AP

Sequential inductive prediction intervals

In this paper we explore the concept of sequential inductive prediction intervals using theory from sequential testing. We furthermore introduce a 3-parameter PAC definition of prediction intervals that allows us via simulation to achieve almost sharp bounds with high probability.

stat.ME

Concentration inequalities for leave-one-out cross validation

In this article we prove that estimator stability is enough to show that leave-one-out cross validation is a sound procedure, by providing concentration bounds in a general framework. In particular, we provide concentration bounds beyond Lipschitz continuity assumptions on the loss or on the estimator. We obtain our results by relying on random variables with distribution satisfying the logarithmic Sobolev inequality, providing us a relatively rich class of distributions. We illustrate our method by considering several interesting examples, including linear regression, kernel density estimation, and stabilized/truncated estimators such as stabilized kernel regression.

math.ST

A Galerkin type method for kinetic Fokker Planck equations based on Hermite expansions

In this paper, we develop a Galerkin-type approximation, with quantitative error estimates, for weak solutions to the Cauchy problem for kinetic Fokker-Planck equations in the domain $(0, T) \times D \times \mathbb{R}^d$, where $D$ is either $\mathbb{T}^d$ or $\mathbb{R}^d$. Our approach is based on a Hermite expansion in the velocity variable only, with a hyperbolic system that appears as the truncation of the Brinkman hierarchy, as well as ideas from $\href{arXiv:1902.04037v2}{AAMN21}$ and additional energy-type estimates that we have developed. We also establish the regularity of the solution based on the regularity of the initial data and the source term.

math.AP

Geometric Characterization of the Eyring-Kramers Formula

In this paper we consider the mean transition time of an over-damped Brownian particle between local minima of a smooth potential. When the minima and saddles are non-degenerate this is in the low noise regime exactly characterized by the so called Eyring-Kramers law and gives the mean transition time as a quantity depending on the curvature of the minima and the saddle. In this paper we find an extension of the Eyring-Kramers law giving an upper bound on the mean transition time when both the minima/saddles are degenerate (flat) while at the same time covering multiple saddles at the same height. Our main contribution is a new sharp characterization of the capacity of two local minimas as a ratio of two geometric quantities, i.e., the smallest separating surface and the geodesic distance.

math.AP

Uncertainty-Aware Body Composition Analysis with Deep Regression Ensembles on UK Biobank MRI

Along with rich health-related metadata, medical images have been acquired for over 40,000 male and female UK Biobank participants, aged 44-82, since 2014. Phenotypes derived from these images, such as measurements of body composition from MRI, can reveal new links between genetics, cardiovascular disease, and metabolic conditions. In this work, six measurements of body composition and adipose tissues were automatically estimated by image-based, deep regression with ResNet50 neural networks from neck-to-knee body MRI. Despite the potential for high speed and accuracy, these networks produce no output segmentations that could indicate the reliability of individual measurements. The presented experiments therefore examine uncertainty quantification with mean-variance regression and ensembling to estimate individual measurement errors and thereby identify potential outliers, anomalies, and other failure cases automatically. In 10-fold cross-validation on data of about 8,500 subjects, mean-variance regression and ensembling showed complementary benefits, reducing the mean absolute error across all predictions by 12%. Both improved the calibration of uncertainties and their ability to identify high prediction errors. With intra-class correlation coefficients (ICC) above 0.97, all targets except the liver fat content yielded relative measurement errors below 5%. Testing on another 1,000 subjects showed consistent performance, and the method was finally deployed for inference to 30,000 subjects with missing reference values. The results indicate that deep regression ensembles could ultimately provide automated, uncertainty-aware measurements of body composition for more than 120,000 UK Biobank neck-to-knee body MRI that are to be acquired within the coming years.

eess.IV

Deep limits and cut-off phenomena for neural networks

We consider dynamical and geometrical aspects of deep learning. For many standard choices of layer maps we display semi-invariant metrics which quantify differences between data or decision functions. This allows us, when considering random layer maps and using non-commutative ergodic theorems, to deduce that certain limits exist when letting the number of layers tend to infinity. We also examine the random initialization of standard networks where we observe a surprising cut-off phenomenon in terms of the number of layers, the depth of the network. This could be a relevant parameter when choosing an appropriate number of layers for a given learning task, or for selecting a good initialization procedure. More generally, we hope that the notions and results in this paper can provide a framework, in particular a geometric one, for a part of the theoretical understanding of deep neural networks.

cs.LG