SearcharxivSearch

arXiv subjects

Max Grieshammer

Publications and source records attributed to Max Grieshammer.

7 recordsLinked to original sources

The Continuous Stochastic Gradient Method: Part II -- Application and Numerics

In this contribution, we present a numerical analysis of the continuous stochastic gradient (CSG) method, including applications from topology optimization and convergence rates. In contrast to standard stochastic gradient optimization schemes, CSG does not discard old gradient samples from previous iterations. Instead, design dependent integration weights are calculated to form a linear combination as an approximation to the true gradient at the current design. As the approximation error vanishes in the course of the iterations, CSG represents a hybrid approach, starting off like a purely stochastic method and behaving like a full gradient scheme in the limit. In this work, the efficiency of CSG is demonstrated for practically relevant applications from topology optimization. These settings are characterized by both, a large number of optimization variables \textit{and} an objective function, whose evaluation requires the numerical computation of multiple integrals concatenated in a nonlinear fashion. Such problems could not be solved by any existing optimization method before. Lastly, with regards to convergence rates, first estimates are provided and confirmed with the help of numerical experiments.

math.OC

The Continuous Stochastic Gradient Method: Part I -- Convergence Theory

In this contribution, we present a full overview of the continuous stochastic gradient (CSG) method, including convergence results, step size rules and algorithmic insights. We consider optimization problems in which the objective function requires some form of integration, e.g., expected values. Since approximating the integration by a fixed quadrature rule can introduce artificial local solutions into the problem while simultaneously raising the computational effort, stochastic optimization schemes have become increasingly popular in such contexts. However, known stochastic gradient type methods are typically limited to expected risk functions and inherently require many iterations. The latter is particularly problematic, if the evaluation of the cost function involves solving multiple state equations, given, e.g., in form of partial differential equations. To overcome these drawbacks, a recent article introduced the CSG method, which reuses old gradient sample information via the calculation of design dependent integration weights to obtain a better approximation to the full gradient. While in the original CSG paper convergence of a subsequence was established for a diminishing step size, here, we provide a complete convergence analysis of CSG for constant step sizes and an Armijo-type line search. Moreover, new methods to obtain the integration weights are presented, extending the application range of CSG to problems involving higher dimensional integrals and distributed data.

math.OC

CSG: A stochastic gradient method for a wide class of optimization problems appearing in a machine learning or data-driven context

A recent article introduced thecontinuous stochastic gradient method (CSG) for the efficient solution of a class of stochastic optimization problems. While the applicability of known stochastic gradient type methods is typically limited to expected risk functions, no such limitation exists for CSG. This advantage stems from the computation of design dependent integration weights, allowing for optimal usage of available information and therefore stronger convergence properties. However, the nature of the formula used for these integration weights essentially limited the practical applicability of this method to problems in which stochasticity enters via a low-dimensional and sufficiently simple probability distribution. In this paper we significantly extend the scope of the CSG method by presenting alternative ways to calculate the integration weights. A full convergence analysis for this new variant of the CSG method is presented and its efficiency is demonstrated in comparison to more classical stochastic gradient methods by means of a number of problem classes relevant to stochastic optimization and machine learning.

math.OC

Measure representation of evolving genealogies

We study evolving genealogies, i.e. processes that take values in the space of (marked) ultra-metric measure spaces and satisfy some sort of "consistency" condition. This condition is based on the observation that the genealogical distance of two individuals who do not have common ancestors up to a time $h$ in the past is completely determined by the genealogical distance of the respective ancestors at that time $h$ in the past. Now the idea is to color all possible ancestors at time $h$ in the past and measure the relative number of their descendants. The resulting collection of measure-valued processes (the construction is possible for all $h$) is called a measure representation. As a main result we give a tightness criterion of evolving genealogies in terms of their measure representation. We then apply our theory to study a finite system scheme for tree-valued interacting Fleming-Viot processes.

math.PR

Family size decomposition of genealogical trees

We study the path of family size decompositions of varying depth of genealogical trees. We prove that this decomposition as a function on (equivalence classes of) ultra-metric measure spaces to the Skorohod space describing the family sizes at different depths is perfect onto its image, i.e. there is a suitable topology such that this map is continuous closed surjective and pre-images of compact sets are compact. We also specify a (dense) subset so that the restriction of the function to this subspace is a homeomorphism. This property allows us to argue that the whole genealogy of a Fleming-Viot process with mutation and selection as well as the genealogy in a Feller branching population can be reconstructed by the genealogical distance of two randomly chosen individuals.

math.PR

Genealogical distance under selection

We study the genealogical distance of two randomly chosen individuals in a population that evolves according to a two type Moran model with mutation and selection. We prove that this distance is stochastically smaller than the corresponding distance in the neutral model, when the population size is large. Moreover, we prove convergence of the genealogical distance under selection to the distance in the neutral case, when the system is in equilibrium and the selection parameter tends to infinity.

math.PR

Partial orders on metric measure spaces

A partial order on the set of metric measure spaces is defined; it generalizes the Lipschitz order of Gromov. We show that our partial order is closed when metric measure spaces are equipped with the Gromov-weak topology and give a new characterization for the Lipschitz order. We will then consider some probabilistic applications. The main importance is given to the study of Fleming-Viot processes with different resampling rates. Besides that application we also consider tree-valued branching processes and two semigroups on metric measure spaces.

math.PR